详细信息

Emotion recognition based on time-scale heterogeneity and hierarchical spatial coupling analysis of multimodal physiological signals  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Emotion recognition based on time-scale heterogeneity and hierarchical spatial coupling analysis of multimodal physiological signals

作者:Xu, Zhangyong[1];Chen, Ning[1];Li, Guangqiang[1];Li, Jing[1];Zhu, Hongqing[1];Zhu, Zhiying[1]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China

年份:2025

卷号:285

外文期刊名:EXPERT SYSTEMS WITH APPLICATIONS

收录:;EI(收录号:20252018410328);WOS:【SCI-EXPANDED(收录号:WOS:001503884900010)】;

基金:Acknowledgement This work was supported by the National Natural Science Foundation of Chinagrant number 61771196, 61872143] . We would like to thank the authors of references (Jiang et al., 2023; Wang et al., 2024a) for answering our questions and the author of reference (Tang et al., 2024) for providing us with the pre-processing features.

语种:英文

外文关键词:Multimodal fusion; Emotion recognition; Time-scale heterogeneity; Hierarchical spatial coupling; Cross-Scale Transformer

摘要:The study on emotion recognition based on the fusion of multimodal physiological signals has attracted much attention. However, the time-scale heterogeneity in each modality and hierarchical spatial coupling patterns contained in each modality or among different modalities, which may affect the fusion effectiveness, are seldom considered. To solve this problem, on the one hand, the Modified Deep-Wise Separable Convolution (MDWSC) layers with different kernel sizes are adopted to extract different time-scale temporal features from each modality, on which Cross-Scale Transformer (CST) is performed to reduce the time-scale heterogeneity among them to achieve effective fusion of them. On the other hand, for each modality, Graph Convolution Network (GCN) with multiple layers is performed on the graph generated by viewing the fused multi-scale temporal feature of each channel as the node feature to grasp multi-level intra-modal coupling features, and then DGCNN is performed on the graph constructed by viewing the aggregated intra-modal coupling feature of each modality in the same level as node feature to grasp the inter-modal coupling feature of corresponding level. Finally, the highest level of intra-modal coupling feature of each modality and the inter-modal coupling features of different levels are fused for emotion prediction. Extensive experimental results on three open datasets demonstrate that the proposed model outperforms State-Of-The-Art (SOTA) fusion baselines in binary and four-class classification under two validation scenarios, and the multi-scale temporal feature extraction and the hierarchical intra-modal and inter-modal coupling features extraction contribute to the performance enhancement of the proposed model.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心