详细信息

Incongruity-aware multimodal physiology signals fusion for emotion recognition  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Incongruity-aware multimodal physiology signals fusion for emotion recognition

作者:Li, Jing[1];Chen, Ning[1];Zhu, Hongqing[1];Li, Guangqiang[1];Xu, Zhangyong[1];Chen, Dingxin[1]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China

年份:2024

卷号:105

外文期刊名:INFORMATION FUSION

收录:;EI(收录号:20240215357979);WOS:【SCI-EXPANDED(收录号:WOS:001152686800001)】;

基金:Acknowledgments This work was supported by the National Natural Science Founda-tion of China [grant number 61771196, 61872143] . The authors would like to thank the anonymous reviewers and the associate editor for their insightful comments that significantly improved the quality of this paper.

语种:英文

外文关键词:Emotion recognition; Multi-modal fusion; Physiological signal incongruity; Cross-Modal Transformer (CMT); Self-Attention Transformer (SAT); Low Rank Fusion (LRF)

摘要:Various physiological signals can reflect the human's emotional states objectively. How to take advantage of the common as well as complementary properties of different physiological signals in representing the emotional states is an interesting problem. Although various models have been constructed to fuse multimodal physiological signals for emotion recognition, the possible incongruity existing among different physiological signals in representing the emotional states and the redundancy resulted from the fusion, which may affect the performance of the fusion schemes seriously, were seldom considered. To this end, a fusion model, which can eliminate the incongruity among different physiological signals and reduce the information redundancy to some extent, is proposed. First, one physiological signal is chosen as the primary modality due to its prominent performance in emotion recognition, and the remaining physiological signals are viewed as the auxiliary modalities. Secondly, the Cross Modal Transformer (CMT) is adopted to optimize the features of the auxiliary modalities by eliminating the incongruity among them, and then Low Rank Fusion (LRF) is performed to eliminate information redundancy caused by fusion. Thirdly, the modified CMT (MCMT) is constructed to enhance the feature of the primary modality by that of each optimized auxiliary modality feature. Fourthly, Self -Attention Transformer (SAT) is performed on the concatenation result of all the enhanced primary modality features to take full advantage of the common as well complementary properties among them in representing the emotional states. Finally, the enhanced primary modality feature and the optimized auxiliary features are fused by concatenation for emotion recognition. Extensive experimental results on DEAP and WESAD datasets demonstrate that (i) The incongruity does exist among different physiological signals, and the CMT-based auxiliary modality feature optimization strategy can eliminate the incongruity prominently; (ii) The emotion prediction accuracy of the primary modality can be enhanced by the auxiliary modality; (iii) All the key modules in the proposed model, CMT, LRF, and MCMT, contribute to the performance enhancement of the proposed model; iv) The proposed model outperforms State -Of -The -Art (SOTA) models in emotion recognition task.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心