详细信息

Exploiting the potentialities of features for speech emotion recognition  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Exploiting the potentialities of features for speech emotion recognition

作者:Li, Dongdong[1];Zhou, Yijun[1];Wang, Zhe[1];Gao, Daqi[1]

机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China

年份:2021

卷号:548

起止页码:328

外文期刊名:INFORMATION SCIENCES

收录:;EI(收录号:20204409423318);WOS:【SCI-EXPANDED(收录号:WOS:000596037700002)】;

基金:This work is supported by Natural Science Foundation of China under Grant No. 61806078,62076094 and 61976091, National Major Scientific and Technological Special Project for "Significant New Drugs Development"under grant no. 2019ZX09201004, Shanghai Science and Technology Program "Distributed and generative few-shot algorithm and theory research" under Grant No.20511100600.

语种:英文

外文关键词:Feature optimization; Feature selection; Speech emotion recognition; Deep learning

摘要:In recent years, studies on speech signals have increasingly paid attention to emotional information. The most challenging aspect in speech emotion recognition (SER) is choosing the optimal speech feature representation. According to the statistical analysis, the roles of each speech feature differ under different emotions, indicating that different features have different abilities in distinguishing emotions. This study proposes an emotional-category based feature weighting (ECFW) method, which aims at finding the prominence of each feature under different emotions and applying this prominence as priori knowledge. Furthermore, previous studies have paid little attention to matching the relationship between speech features and models. This study argues that different combinations of models and features result in large differences in the performance of SER, which are evaluated by several experiments. Features must be modeled with appropriate approaches to extract the most valuable information for emotional representation. Then, the best combinations of features and models are selected to test our method. The method is applied on three commonly used speech emotion databases, IEMOCAP, MASC, and EMO-DB. The results show that ECFW significantly improves the performance of SER tasks. (C) 2020 Elsevier Inc. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心