详细信息
Deep multi-modal fusion transformer for emotion recognition ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Deep multi-modal fusion transformer for emotion recognition
作者:Zhang, Qian[1,2];Liu, Yifan[2];Zhu, Biaokai[3];Han, Xun[4];Zhang, Ruilin[2];Xiao, Jun[2];Wang, Zhe[1,2]
机构:[1]Minist Educ, Key Lab Smart Mfg Energy Chem Proc, Shanghai, Peoples R China;[2]East China Univ Sci & Technol, Shanghai, Peoples R China;[3]Shanxi Police Coll, Taiyuan, Peoples R China;[4]Sichuan Police Coll, Intelligent Policing Key Lab Sichuan Prov, Luzhou, Peoples R China
年份:2026
卷号:168
外文期刊名:ENGINEERING APPLICATIONS OF ARTIFICIAL INTELLIGENCE
收录:;EI(收录号:20260620026028);WOS:【SCI-EXPANDED(收录号:WOS:001684134000001)】;
基金:The work of Zhe Wang is supported by Natural Science Foundation of China (No. 62476087). The work of Qian Zhang and Xun Han is partially supported by Intelligent Policing Key Laboratory of Sichuan Province (No. ZNJW2024KFQN006). The work of Biaokai Zhu is supported by National Natural Science Foundation of China (No. 62306207), Program for the Young Academic Leaders of Higher Learning Institutions of Shanxi (No. 2024Q042).
语种:英文
外文关键词:Emotion recognition; Multi-modal fusion; Cross-attention; Transformers
摘要:Multi-modal emotion recognition methods usually integrate peripheral and physiological information to extract complementary features. However, current multi-modal methods face shortcomings in spatio-temporal dependency modeling and feature complementary feature extraction across modalities, leading to limitations in accuracy and robustness for emotion recognition. To address these issues, this paper proposes a multi-modal emotion recognition network based on cross-modal transformer fusion to enhance the collaborative processing ability of electroencephalogram (EEG) and facial expression data. Specifically, we construct a fine and coarse combined transformer encoder for multi-level spatiotemporal feature extraction of EEG signals, and introduce a multi-modal cross-attention fusion transformer to achieve deep fusion of EEG, facial expression features and joint features, capturing dynamic relationships between and within modalities. Experimental results in two public datasets show that this method outperforms existing methods in accuracy and robustness, achieving a deep fusion of emotional dynamic and complex characteristics.
参考文献:
正在载入数据...
