详细信息
Multi-level Feature Joint Learning Methods for Emotional Speaker Recognition ( EI收录)
文献类型:期刊文献
英文题名:Multi-level Feature Joint Learning Methods for Emotional Speaker Recognition
作者:Zeng, Zhongliang[1]; Li, Dongdong[1]; Wang, Zhe[1]; Yang, Hai[1]
机构:[1] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China
年份:2023
卷号:2023-June
外文期刊名:Proceedings of the International Joint Conference on Neural Networks
收录:EI(收录号:20233614678786)
语种:英文
外文关键词:Emotion Recognition - Speech recognition
摘要:In the real scene, changes in speaker features caused by different emotional states have a great impact on the performance of speaker recognition. To improve the robustness of the speaker recognition system, the existing emotional speaker recognition technologies tend to cascade different models, ignoring the frame-level acoustic features and the segment-level discourse habits feature. To this end, we combine frame- and segment-level features in different ways to build a robust recognition system for emotional speakers. The frame-level features and segment-level features are jointly learned to retain emotional information and speaker information. Four joint learning methods, namely, Joint in series, Joint in Parallel, Joint under the guidance, and Joint with Original Feature, are discussed to explore the correlations between fragment-level features and frame-level features. The experimental results illustrate that the speaker feature will change greatly in different emotional states. Compared with the accuracy of 90.95% by x-vector, the proposed methods of Joint in parallel and Joint with Original Features can achieve the accuracy of 95.06% and 94.67% respectively for emotional speaker recognition in the experiment on Mandarin Affective Speech Corpus (MASC). Our findings provide a novel aspect to improve speaker recognition robustness. ? 2023 IEEE.
参考文献:
正在载入数据...
