详细信息
Dynamic facial expression recognition based on spatial key-points optimized region feature fusion and temporal self-attention ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Dynamic facial expression recognition based on spatial key-points optimized region feature fusion and temporal self-attention
作者:Huang, Zhiwei[1];Zhu, Yu[1];Li, Hangyu[1];Yang, Dawei[2,3]
机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Fudan Univ, Zhongshan Hosp, Dept Pulm & Crit Care Med, Shanghai 200032, Peoples R China;[3]Shanghai Engn Res Ctr Internet Things Resp Med, Shanghai 200032, Peoples R China
年份:2024
卷号:133
外文期刊名:ENGINEERING APPLICATIONS OF ARTIFICIAL INTELLIGENCE
收录:;EI(收录号:20242016105003);WOS:【SCI-EXPANDED(收录号:WOS:001293703100001)】;
基金:All authors discuss the results and contribute to the final manuscript. This work is supported in part by the National Natural Science Foundation of China under Grant 82170110, and the Science and Technology Commission of Shanghai Municipality under Grant 20DZ22544000, 21DZ2200600, 20DZ2261200. Fujian Province Department of Science and Technology (2022D014) .
语种:英文
外文关键词:Dynamic facial expression recognition; Spatial feature fusion; Graph convolution network; Self-attention
摘要:Dynamic facial expression recognition (DFER) is of great significance in promoting empathetic machines and metaverse technology. However, dynamic facial expression recognition (DFER) in the wild remains a challenging task, often constrained by complex lighting changes, frequent key-points occlusion, uncertain emotional peaks and severe imbalanced dataset categories. To tackle these problems, this paper presents a depth neural network model based on spatial key-points optimized region feature fusion and temporal self- attention. The method includes three parts: spatial feature extraction module, temporal feature extraction module and region feature fusion module. The intra-frame spatial feature extraction module is composed of the key-points graph convolution network (GCN) and a convolution network (CNN) branch to obtain the global and local feature vectors. The newly proposed region fusion strategy based on face spatial structure is used to obtain the spatial fusion feature of each frame. The inter-frame temporal feature extraction module uses multi-head self-attention model to obtain the temporal information of inter-frames. The experimental results show that our method achieves accuracy of 68.73%, 55.00%, 47.80%, and 47.44% on the DFEW, AFEW, FERV39k, and MAFW datasets. Ablation experiments showed that the GCN module, fusion module, and temporal module improved the accuracy on DFEW by 0.68%, 1.66%, and 3.25%, respectively. The method also achieves competitive results in terms of parameter quantity and inference speed, which demonstrates the effectiveness of the proposed method.
参考文献:
正在载入数据...
