详细信息
Emotion embedding framework with emotional self-attention mechanism for speaker recognition ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Emotion embedding framework with emotional self-attention mechanism for speaker recognition
作者:Li, Dongdong[1];Yang, Zhuo[1];Liu, Jinlin[1];Yang, Hai[1];Wang, Zhe[1]
机构:[1]East China Univ Sci & Technol, Dept Informat Sci & Engn, Shanghai 200237, Peoples R China
年份:2024
卷号:238
外文期刊名:EXPERT SYSTEMS WITH APPLICATIONS
收录:;EI(收录号:20234515010026);WOS:【SCI-EXPANDED(收录号:WOS:001107337400001)】;
基金:This work is supported by Natural Science Foundation of China under Grant No. 62276098, No. 62076094, Shanghai Science and Technology Program "Federated based cross-domain and cross-task incremental learning" under Grant No. 21511100800.
语种:英文
外文关键词:Speaker recognition; Emotional states; Emotion embedding; Emotional self-attention
摘要:The emotional states of speech have a great impact on the efficiency of speaker recognition (SR) system. Many researchers focus on how to map speech with different emotions to an emotion invariant embedding, which reduces the diversity of data. This paper proposes a new emotion embedding framework with self attention mechanism for speaker recognition. First, several deep neural networks (DNNs) are trained to classify speakers in different emotional states as emotion embedding extractors during development phase. Then at enrollment stage, these pre-trained models are used to extend emotion embeddings from neutral speech. In order to make the final speaker embedding more representative, the classification model is trained with self attention mechanism in emotion dimension, so that the framework can automatically annotate the weights of the emotion embeddings. Experiments were carried out on both Mandarin Affective Speech Corpus (MASC) and Crowd-Sourced Emotional Multimodal Actors Dataset (CREMA-D). The results show the proposed method achieves the best of Identification Rate (IR) and Equal Error Rate (EER) which are 59.14%, 15.79% on MASC and 75.98%, 8.14% on CREMA-D compared with state-of-the-art methods. In addition, the cross-database experiments also further demonstrate the practicability of the method in real scenes.
参考文献:
正在载入数据...
