详细信息

Multi-level Feature Joint Learning Methods for Emotional Speaker Recognition  ( EI收录)  

文献类型:期刊文献

英文题名:Multi-level Feature Joint Learning Methods for Emotional Speaker Recognition

作者:Zeng, Zhongliang[1]; Li, Dongdong[1]; Wang, Zhe[1]; Yang, Hai[1]

机构:[1] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China

年份:2023

卷号:2023-June

外文期刊名:Proceedings of the International Joint Conference on Neural Networks

收录:EI(收录号:20233614678786)

语种:英文

外文关键词:Emotion Recognition - Speech recognition

摘要:In the real scene, changes in speaker features caused by different emotional states have a great impact on the performance of speaker recognition. To improve the robustness of the speaker recognition system, the existing emotional speaker recognition technologies tend to cascade different models, ignoring the frame-level acoustic features and the segment-level discourse habits feature. To this end, we combine frame- and segment-level features in different ways to build a robust recognition system for emotional speakers. The frame-level features and segment-level features are jointly learned to retain emotional information and speaker information. Four joint learning methods, namely, Joint in series, Joint in Parallel, Joint under the guidance, and Joint with Original Feature, are discussed to explore the correlations between fragment-level features and frame-level features. The experimental results illustrate that the speaker feature will change greatly in different emotional states. Compared with the accuracy of 90.95% by x-vector, the proposed methods of Joint in parallel and Joint with Original Features can achieve the accuracy of 95.06% and 94.67% respectively for emotional speaker recognition in the experiment on Mandarin Affective Speech Corpus (MASC). Our findings provide a novel aspect to improve speaker recognition robustness. ? 2023 IEEE.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心