详细信息

Heterogeneity-aware multi-modal physiological signal fusion strategy based on combined contrastive learning for emotion recognition  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Heterogeneity-aware multi-modal physiological signal fusion strategy based on combined contrastive learning for emotion recognition

作者:Tian, Yuan[1];Li, Jing[1];Chen, Ning[1];Li, Guangqiang[1];Xu, Zhangyong[1];Zhu, Hongqing[1];Li, Yu[1];Zhu, Zhiying[1]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China

年份:2026

卷号:200

外文期刊名:NEURAL NETWORKS

收录:;EI(收录号:20261520499263);WOS:【SCI-EXPANDED(收录号:WOS:001717170600001)】;

基金:This work was supported by the National Natural Science Foundation of China [grant number 61771196, 61872143] .

语种:英文

外文关键词:Multi-modal emotion recognition; Physiological signals; Heterogeneity; Contrastive learning

摘要:Due to the inherent non-stationary nature and significant cross-subject divergence in physiological signals, there exist prominent heterogeneity among different modalities, different channels, and different temporal patches, which will influence the fusion effectiveness of multimodal physiological signals greatly. To mitigate above heterogeneities simultaneously and enhance multimodal emotion recognition performance, a combined cross-modal contrastive learning strategy is proposed in this paper. First, a Graph Attention Network (GAT) based learnable view augmentation is introduced to simulate the variations introduced by the non-stationary nature and cross-subject divergence. Next, the temporal contrastive learning is performed between the current temporal patch and its previous temporal patches in the augmented view to mitigate the heterogeneity among different temporal patches. Then, the cross-channel contrastive learning is performed within each view and between different views to reduce both the cross-modal and cross-channel heterogeneities. Extensive experimental results under both cross-trial and cross-subject scenarios on DEAP, DREAMER, and PhyMER datasets demonstrate that: i) The proposed model outperforms state-of-the-art (SOTA) multimodal fusion models. ii) The learnable view augmentation, the temporal contrastive learning strategy, and the spatial contrastive learning strategy contribute to the performance enhancement of the proposed model. iii) The proposed model can take full advantage of the complementarities among different modalities in representing emotional states.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心