详细信息
基于图神经网络多模态融合的语音情感识别模型
Speech emotion recognition based on multi-modal fusion of graph neural network
文献类型:期刊文献
中文题名:基于图神经网络多模态融合的语音情感识别模型
英文题名:Speech emotion recognition based on multi-modal fusion of graph neural network
作者:李紫荆[1];陈宁[1]
机构:[1]华东理工大学信息科学与工程学院,上海200237
年份:2023
卷号:40
期号:8
起止页码:2286
中文期刊名:计算机应用研究
外文期刊名:Application Research of Computers
收录:CSTPCD;;北大核心:【北大核心2020】;CSCD:【CSCD_E2023_2024】;
基金:国家自然科学基金资助项目(61771196)。
语种:中文
中文关键词:语音情感识别;多模态特征;图神经网络;图增强
外文关键词:speech emotion recognition;multi-modal feature;graph neural network;graph augmentation
摘要:目前,基于多模态融合的语音情感识别模型普遍存在无法充分利用多模态特征之间的共性和互补性、无法借助样本特征间的拓扑结构特性对样本特征进行有效地优化和聚合,以及模型复杂度过高的问题。为此,引入图神经网络,一方面在特征优化阶段,将经过图神经网络优化后的文本特征作为共享表示重构基于声学特征的邻接矩阵,使得在声学特征的拓扑结构特性中包含文本信息,达到多模态特征的融合效果;另一方面在标签预测阶段,借助图神经网络充分聚合当前节点的邻接节点所包含的相似性信息对当前节点特征进行全局优化,以提升情感识别准确率。同时为防止图神经网络训练过程中可能出现的过平滑问题,在图神经网络训练前先进行图增强处理。在公开数据集IEMOCAP和RAVDESS上的实验结果表明,所提出的模型取得了比基线模型更高的识别准确率和更低的模型复杂度,并且模型各个组成部分均对模型性能提升有所贡献。
At present,speech emotion recognition models based on multi-modal fusion generally suffer from the inability to make full use of the commonality and complementarity between multimodal features,the inability to effectively optimize and aggregate sample features by using the topological structure characteristics between sample features,and the high complexity of existing models.Therefore,this paper introduced graph neural network.On the one hand,in the feature optimization stage,it used the text features optimized by the graph neural network as a shared representation to reconstruct the adjacency matrix based on acoustic features,so that the topological structure characteristics of the acoustic features contained text information,thus achieving multi-modal fusion.On the other hand,in the label prediction stage,it used the graph neural network to fully aggregate the similarity information contained in the adjacent nodes of the current node to optimize the characteristics of the current node globally to improve the accuracy of emotion recognition.At the same time,in order to prevent the over-smoothing problem that might occur during the training of the graph neural network,it performed graph augmentation before the graph neural network training.The experimental results on the public datasets IEMOCAP and RAVDESS show that the proposed model achieves higher recognition accuracy and lower model complexity than the baseline models,and each component of the model contributes to the improvement of model performance.
参考文献:
正在载入数据...
