详细信息

Speech emotion recognition using recurrent neural networks with directional self-attention  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Speech emotion recognition using recurrent neural networks with directional self-attention

作者:Li, Dongdong[1,2];Liu, Jinlin[1];Yang, Zhuo[1];Sun, Linyu[1];Wang, Zhe[1]

机构:[1]East China Univ Sci & Technol, Dept Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Soochow Univ, Prov Key Lab Comp Informat Proc Technol, Suzhou 215006, Peoples R China

年份:2021

卷号:173

外文期刊名:EXPERT SYSTEMS WITH APPLICATIONS

收录:;EI(收录号:20211010026432);WOS:【SCI-EXPANDED(收录号:WOS:000636782600004)】;

基金:This work is supported by Natural Science Foundation of China under Grant No. 61806078, No. 62076094, National Major Scientific and Technological Special Project for "Significant New Drugs Development" under Grant No. 2019ZX09201004, Shanghai Science and Technology Program "Distributed and generative few-shot algorithm and theory research" under Grant No. 20511100600.

语种:英文

外文关键词:Speech emotion recognition; Bi-directional long-short term memory with directional self-attention; Self-attention; Autocorrelation

摘要:As an important branch of affective computing, Speech Emotion Recognition (SER) plays a vital role in human?computer interaction. In order to mine the relevance of signals in audios an increase the diversity of information, Bi-directional Long-Short Term Memory with Directional Self-Attention (BLSTM-DSA) is proposed in this paper. Long Short-Term Memory (LSTM) can learn long-term dependencies from learned local features. Moreover, Bi-directional Long-Short Term Memory (BLSTM) can make the structure more robust by direction mechanism because that the directional analysis can better recognize the hidden emotions in sentence. At the same time, autocorrelation of speech frames can be used to deal with the lack of information, so that SelfAttention mechanism is introduced into SER. The attention weight of each frame is calculated with the output of the forward and backward LSTM respectively rather than calculated after adding them together. Thus, the algorithm can automatically annotate the weights of speech frames to correctly select frames with emotional information in temporal network. When evaluate it on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) database and Berlin database of emotional speech (EMO-DB), the BLSTM-DSA demonstrates satisfactory performance on the task of speech emotion recognition. Especially in emotion recognizing of happiness and anger, BLSTM-DSA achieves the highest recognition accuracies.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心