详细信息
Heterogeneity-aware multi-modal physiological signal fusion strategy based on combined contrastive learning for emotion recognition ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Heterogeneity-aware multi-modal physiological signal fusion strategy based on combined contrastive learning for emotion recognition
作者:Tian, Yuan[1];Li, Jing[1];Chen, Ning[1];Li, Guangqiang[1];Xu, Zhangyong[1];Zhu, Hongqing[1];Li, Yu[1];Zhu, Zhiying[1]
机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China
年份:2026
卷号:200
外文期刊名:NEURAL NETWORKS
收录:;EI(收录号:20261520499263);WOS:【SCI-EXPANDED(收录号:WOS:001717170600001)】;
基金:This work was supported by the National Natural Science Foundation of China [grant number 61771196, 61872143] .
语种:英文
外文关键词:Multi-modal emotion recognition; Physiological signals; Heterogeneity; Contrastive learning
摘要:Due to the inherent non-stationary nature and significant cross-subject divergence in physiological signals, there exist prominent heterogeneity among different modalities, different channels, and different temporal patches, which will influence the fusion effectiveness of multimodal physiological signals greatly. To mitigate above heterogeneities simultaneously and enhance multimodal emotion recognition performance, a combined cross-modal contrastive learning strategy is proposed in this paper. First, a Graph Attention Network (GAT) based learnable view augmentation is introduced to simulate the variations introduced by the non-stationary nature and cross-subject divergence. Next, the temporal contrastive learning is performed between the current temporal patch and its previous temporal patches in the augmented view to mitigate the heterogeneity among different temporal patches. Then, the cross-channel contrastive learning is performed within each view and between different views to reduce both the cross-modal and cross-channel heterogeneities. Extensive experimental results under both cross-trial and cross-subject scenarios on DEAP, DREAMER, and PhyMER datasets demonstrate that: i) The proposed model outperforms state-of-the-art (SOTA) multimodal fusion models. ii) The learnable view augmentation, the temporal contrastive learning strategy, and the spatial contrastive learning strategy contribute to the performance enhancement of the proposed model. iii) The proposed model can take full advantage of the complementarities among different modalities in representing emotional states.
参考文献:
正在载入数据...
