详细信息

Multi-type features separating fusion learning for Speech Emotion Recognition  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Multi-type features separating fusion learning for Speech Emotion Recognition

作者:Xu, Xinlei[1,2];Li, Dongdong[2];Zhou, Yijun[2];Wang, Zhe[1,2]

机构:[1]East China Univ Sci Technol, Key Lab Smart Mfg Energy Chem Proc, Minist Educ, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China

年份:2022

卷号:130

外文期刊名:APPLIED SOFT COMPUTING

收录:;EI(收录号:20224012846199);WOS:【SCI-EXPANDED(收录号:WOS:000907771700005)】;

基金:This work is supported by National Natural Science Foundation of China under Grant No. 62276098 and No. 62076094, Shanghai Science and Technology Program, China "Federated based crossdomain and cross-task incremental learning" under Grant No. 21511100800, Shanghai Science and Technology Program, China "Distributed and generative few-shot algorithm and theory research" under Grant No. 20511100600, Chinese Defense Program of Science and Technology under Grant No. 2021-JCJQ-JJ-0041, China Aerospace Science and Technology Corporation IndustryUniversity-Research Cooperation Foundation of the Eighth Research Institute under Grant No. SAST2021-007.

语种:英文

外文关键词:Speech emotion recognition; Hybrid feature selection; Feature-level fusion; Speaker-independent

摘要:Speech Emotion Recognition (SER) is a challengeable task to improve human-computer interaction. Speech data have different representations, and choosing the appropriate features to express the emotion behind the speech is difficult. The human brain can comprehensively judge the same thing in different dimensional representations to obtain the final result. Inspired by this, we believe that it is reasonable to have complementary advantages between different representations of speech data. Therefore, a Hybrid Deep Learning with Multi-type features Model (HD-MFM) is proposed to integrate the acoustic, temporal and image information of speech. Specifically, we utilize Convolutional Neural Network (CNN) to extract image information from the spectrogram of speech. Deep Neural Network (DNN) is used for extracting the acoustic information from the statistic features of speech. Then, Long Short-Term Memory (LSTM) is chosen to extract the temporal information from the Mel-Frequency Cepstral Coefficients (MFCC) of speech. Finally, three different types of speech features are concatenated together to get a richer emotion representation with better discriminative property. Considering that different fusion strategies affect the relationship between features, we consider two fusion strategies in this paper named separating and merging. To evaluate the feasibility and effectiveness of the proposed HD-MFM, we perform extensive experiments on EMO-DB and IEMOCAP of SER. The experimental results show that the separating method has more significant advantages in feature complementarity. The proposed HD-MFM obtains 91.25% and 72.02% results on EMO-DB and IEMOCAP. The obtained results indicate the proposed HD-MFM can make full use of the effective complementary feature representations by separating strategy to further enhance the performance of SER. (c) 2022 Elsevier B.V. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心