详细信息
Speech emotion recognition using recurrent neural networks with directional self-attention ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Speech emotion recognition using recurrent neural networks with directional self-attention
作者:Li, Dongdong[1,2];Liu, Jinlin[1];Yang, Zhuo[1];Sun, Linyu[1];Wang, Zhe[1]
机构:[1]East China Univ Sci & Technol, Dept Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Soochow Univ, Prov Key Lab Comp Informat Proc Technol, Suzhou 215006, Peoples R China
年份:2021
卷号:173
外文期刊名:EXPERT SYSTEMS WITH APPLICATIONS
收录:;EI(收录号:20211010026432);WOS:【SCI-EXPANDED(收录号:WOS:000636782600004)】;
基金:This work is supported by Natural Science Foundation of China under Grant No. 61806078, No. 62076094, National Major Scientific and Technological Special Project for "Significant New Drugs Development" under Grant No. 2019ZX09201004, Shanghai Science and Technology Program "Distributed and generative few-shot algorithm and theory research" under Grant No. 20511100600.
语种:英文
外文关键词:Speech emotion recognition; Bi-directional long-short term memory with directional self-attention; Self-attention; Autocorrelation
摘要:As an important branch of affective computing, Speech Emotion Recognition (SER) plays a vital role in human?computer interaction. In order to mine the relevance of signals in audios an increase the diversity of information, Bi-directional Long-Short Term Memory with Directional Self-Attention (BLSTM-DSA) is proposed in this paper. Long Short-Term Memory (LSTM) can learn long-term dependencies from learned local features. Moreover, Bi-directional Long-Short Term Memory (BLSTM) can make the structure more robust by direction mechanism because that the directional analysis can better recognize the hidden emotions in sentence. At the same time, autocorrelation of speech frames can be used to deal with the lack of information, so that SelfAttention mechanism is introduced into SER. The attention weight of each frame is calculated with the output of the forward and backward LSTM respectively rather than calculated after adding them together. Thus, the algorithm can automatically annotate the weights of speech frames to correctly select frames with emotional information in temporal network. When evaluate it on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) database and Berlin database of emotional speech (EMO-DB), the BLSTM-DSA demonstrates satisfactory performance on the task of speech emotion recognition. Especially in emotion recognizing of happiness and anger, BLSTM-DSA achieves the highest recognition accuracies.
参考文献:
正在载入数据...
