详细信息

Mixed Entropy Down-Sampling based Ensemble Learning for Speech Emotion Recognition  ( EI收录)  

文献类型:期刊文献

英文题名:Mixed Entropy Down-Sampling based Ensemble Learning for Speech Emotion Recognition

作者:Xuan, Zhengji[1]; Li, Dongdong[1]; Wang, Zhe[1]; Yang, Hai[1]

机构:[1] East China University of Science and Technology, Department of Computer Science and Engineering, Shanghai, China

年份:2023

卷号:2023-June

外文期刊名:Proceedings of the International Joint Conference on Neural Networks

收录:EI(收录号:20233614679001)

语种:英文

外文关键词:Adaptive boosting - Convolutional neural networks - Deep learning - Emotion Recognition - Iterative methods - Neural network models - Speech recognition

摘要:The strength of emotion at different positions in a speech is strong or weak, and the weak parts with unclear emotions will bring noise to the model. We propose a boosting ensemble learning method based on mixed entropy down-sampling to effectively select emotionally salient segments to improve the classifier's performance. An independent Convolutional Neural Network (CNN) model is trained in each iteration of ensemble learning. These CNN models form an ensemble classifier, which improves the generalization ability by synthesizing all the learned results of down-sampling, making emotion recognition more accurate. We also introduce the concept of Mixed Information Entropy (MIE), which consists of Emotional Certainty Entropy (ECE) and Structural Distribution Entropy (SDE). ECE measures the emotional confusion of segments, while SDE measures the stability of segments in deep feature space. During the iteration, the deep features are obtained from the last fully connected layer of the model and down-sampled according to the weighted sum of confidence and MIE. The selected segments with stronger emotions are used for the next iteration. Our method is 3.77% higher on WA and 2.37% higher on UA than the naive CNN model on the IEMOCAP dataset. ? 2023 IEEE.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心