详细信息
文献类型:期刊文献
中文题名:基于相似度融合和GCN的音频分类模型
英文题名:AUDIO CLASSIFICATION MODEL BASED ON SIMILARITY FUSION AND GCN
作者:何敏捷[1];陈宁[1];林家骏[1]
机构:[1]华东理工大学信息科学与工程学院,上海200237
年份:2025
卷号:42
期号:5
起止页码:116
中文期刊名:计算机应用与软件
外文期刊名:Computer Applications and Software
收录:;北大核心:【北大核心2023】;
基金:国家自然科学基金面上项目(61771196)。
语种:中文
中文关键词:特征融合;图卷积网络;相似度网络融合;深度学习;音频分类
外文关键词:Feature fusion;Graph convolutional network;Similarity network fusion;Deep learning;Audio classification
摘要:为了充分利用样本间基于不同音频特征的相似度表示的拓扑结构特性的互补性,提出一种基于相似度融合和GCN的音频分类模型。分别利用基于CNN14和DenseNet的网络提取输入音频的特征,并进行音频类别的预测;利用相似度网络融合模型对基于以上两个网络获得的预测标签向量的相似度进行非线性融合;分别用DenseNet提取的特征和融合相似度网络对GCN的节点特征和邻接矩阵进行初始化,通过GCN进行节点特征优化以提升音频分类准确性。实验结果表明,在不同的音频分类任务中,该模型相比于基线模型取得了更高的分类准确率,且基于SNF的相似度融合模块和基于GCN的分类模块均对模型性能的提升有贡献。
In order to take full advantage of the complementarity of topological features represented by similarity of different audio features between samples,an audio classification model based on similarity fusion and GCN is proposed.The network based on CNN14 and DenseNet was used to extract the features of input audio,and the audio class prediction was conducted.The similarity network fusion(SNF)model was used to make nonlinear fusion of the mutual similarity networks,which were based on the similarity of the predicted tag vectors obtained from the above two networks.The features extracted by DenseNet were used to initialize the node features of Graph Convolutional Networks(GCN),and the fusion similarity network was used to initialize the adjacency matrix.The node features optimization by GCN could improve audio classification accuracy significantly.Experimental results show that the proposed model is superior to the baseline model in different audio classification tasks.The similarity fusion module based on SNF and the classification module based on GCN both contribute to the improvement of model performance.
参考文献:
正在载入数据...
