详细信息

BLSTM and CNN Stacking Architecture for Speech Emotion Recognition  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:BLSTM and CNN Stacking Architecture for Speech Emotion Recognition

作者:Li, Dongdong[1,2,3];Sun, Linyu[2];Xu, Xinlei[2];Wang, Zhe[1,2];Zhang, Jing[2];Du, Wenli[1]

机构:[1]East China Univ Sci & Technol, Key Lab Adv Control & Optimizat Chem Proc, Minist Educ, Shanghai 200237, Peoples R China;[2]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China;[3]Soochow Univ, Prov Key Lab Comp Informat Proc Technol, Suzhou 215006, Peoples R China

年份:2021

卷号:53

期号:6

起止页码:4097

外文期刊名:NEURAL PROCESSING LETTERS

收录:;EI(收录号:20213310779194);WOS:【SCI-EXPANDED(收录号:WOS:000683234100002)】;

基金:This work is supported by Natural Science Foundation of China under Grant No. 61806078, No.62076094, No. 61976091, Shanghai Science and Technology Program "Distributed and generative few-shot algorithm and theory research" under Grant No. 20511100600.

语种:英文

外文关键词:Speech emotion recognition; Convolutional neural network; Bidirectional long short term memory; Stacking

摘要:Speech Emotion Recognition (SER) is a huge challenge for distinguishing and interpreting the sentiments carried in speech. Fortunately, deep learning is proved to have great ability to deal with acoustic features. For instance, Bidirectional Long Short Term Memory (BLSTM) has an advantage of solving time series acoustic features and Convolutional Neural Network (CNN) can discover the local structure among different features. This paper proposed the BLSTM and CNN Stacking Architecture (BCSA) to enhance the ability to recognition emotions. In order to match the input formats of BLSTM and CNN, slicing feature matrices is necessary. For utilizing the different roles of the BLSTM and CNN, the Stacking is employed to integrate the BLSTM and CNN. In detail, taking into account overfitting problem, the estimates of probabilistic quantities from BLSTM and CNN are combined as new data using K-fold cross validation. Finally, based on the Stacking models, the logistic regression is used to recognize emotions effectively by fitting the new data. The experiment results demonstrate that the performance of proposed architecture is better than that of single model. Furthermore, compared with the state-of-the-art model on SER in our knowledge, the proposed method BCSA may be more suitable for SER by integrating time series acoustic features and the local structure among different features.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心