详细信息

Hierarchical dynamic pattern analysis and adaptive fusion of multimodal physiological signals for emotion recognition  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Hierarchical dynamic pattern analysis and adaptive fusion of multimodal physiological signals for emotion recognition

作者:Xu, Zhangyong[1];Chen, Ning[1];Li, Guangqiang[1];Tian, Yuan[1];Zhu, Hongqing[1];Zhu, Zhiying[1]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China

年份:2026

卷号:63

期号:6

外文期刊名:INFORMATION PROCESSING & MANAGEMENT

收录:;EI(收录号:20261320365906);WOS:【SSCI(收录号:WOS:001718182500001),SCI-EXPANDED(收录号:WOS:001718182500001)】;

语种:英文

外文关键词:Multimodal physiological signal; Emotion recognition; Dynamic pattern analysis; Adaptive fusion

摘要:The inherent dynamic characteristics that exist in the temporal and spatial domains of different physiological signals, as well as in cross-modal collaborative relationships among them, are highly related to emotional states. However, these properties are overlooked in the previously proposed multimodal emotion recognition models. To address this limitation, a new fusion model that effectively combines the above emotion-related neuroscience prior knowledge is proposed. First, a temporal progressive pattern extraction strategy is introduced by segmenting signals into multiple phases, extracting differential entropy features from each phase, and adopting the modified depthwise separable convolution to capture the temporal progressive pattern across phases. Second, a spatial asymmetric coupling pattern extraction method based on the multi-layer graph attention network is designed to capture multi-scale spatial asymmetric coupling features. Finally, an adaptive fusion strategy based on the cross-modal Transformer is developed to adaptively fuse emotion-related components in modality-consistent information and those in modality-specific information. Experimental results on DEAP, DREAMER, and PhyMER indicate that our model consistently outperforms state-of-the-art baselines, achieving an accuracy improvement of 0.99%-14.46% (0.82%-6.77%) and F1 score improvement of 0.48%-16.23% (0.17%-7.41%) under subject-dependent (subject-independent) scenarios, in binary and four-class classification tasks with equivalent or even smaller computational complexity. Moreover, all the key modules in the above-designed methods contribute to the performance enhancement. Our findings indicate that extracting the dynamic patterns existing in the temporal and spatial domains of each physiological signal, and those reflected in the cross-modal collaborative relationships, enhances the emotion representation ability and interpretability of the proposed model.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心