详细信息

Emotion embedding framework with emotional self-attention mechanism for speaker recognition  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Emotion embedding framework with emotional self-attention mechanism for speaker recognition

作者:Li, Dongdong[1];Yang, Zhuo[1];Liu, Jinlin[1];Yang, Hai[1];Wang, Zhe[1]

机构:[1]East China Univ Sci & Technol, Dept Informat Sci & Engn, Shanghai 200237, Peoples R China

年份:2024

卷号:238

外文期刊名:EXPERT SYSTEMS WITH APPLICATIONS

收录:;EI(收录号:20234515010026);WOS:【SCI-EXPANDED(收录号:WOS:001107337400001)】;

基金:This work is supported by Natural Science Foundation of China under Grant No. 62276098, No. 62076094, Shanghai Science and Technology Program "Federated based cross-domain and cross-task incremental learning" under Grant No. 21511100800.

语种:英文

外文关键词:Speaker recognition; Emotional states; Emotion embedding; Emotional self-attention

摘要:The emotional states of speech have a great impact on the efficiency of speaker recognition (SR) system. Many researchers focus on how to map speech with different emotions to an emotion invariant embedding, which reduces the diversity of data. This paper proposes a new emotion embedding framework with self attention mechanism for speaker recognition. First, several deep neural networks (DNNs) are trained to classify speakers in different emotional states as emotion embedding extractors during development phase. Then at enrollment stage, these pre-trained models are used to extend emotion embeddings from neutral speech. In order to make the final speaker embedding more representative, the classification model is trained with self attention mechanism in emotion dimension, so that the framework can automatically annotate the weights of the emotion embeddings. Experiments were carried out on both Mandarin Affective Speech Corpus (MASC) and Crowd-Sourced Emotional Multimodal Actors Dataset (CREMA-D). The results show the proposed method achieves the best of Identification Rate (IR) and Equal Error Rate (EER) which are 59.14%, 15.79% on MASC and 75.98%, 8.14% on CREMA-D compared with state-of-the-art methods. In addition, the cross-database experiments also further demonstrate the practicability of the method in real scenes.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心