详细信息

Saccade inspired Attentive Visual Patch Transformer for image sentiment analysis  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Saccade inspired Attentive Visual Patch Transformer for image sentiment analysis

作者:Zhang, Jing[1];Zhu, Jixiang[1];Sun, Han[1];Zhang, Xinzhou[1];Liu, Jiangpei[1]

机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai, Peoples R China

年份:2025

卷号:174

外文期刊名:APPLIED SOFT COMPUTING

收录:;EI(收录号:20251118054492);WOS:【SCI-EXPANDED(收录号:WOS:001450613600001)】;

基金:This study was supported by the Natural Science Foundation Shanghai, China "Research on image sentiment analysis and expression based on human vision and cognitive psychology" (22ZR1418400) .

语种:英文

外文关键词:Image sentiment analysis; Attentive visual patch transformer; Visual attention shift; Saccade mechanism

摘要:The generation of image-evoked emotion is usually regarded as a transient process in the image sentiment analysis. However, according to the saccade mechanism of the human visual system, the evoked emotion generated during the saccade process changes over time and attention. Based on above analysis, we propose an Attentive Visual Patch Transformer (AVPT), using visual attention sequence to represent the sentiment context of images and predict the possible distribution of sentiment. In AVPT, the spatial structure in the form of patches are reconstructed and reorganized by visual attention shift sequentially. Simultaneously, the temporal characteristics of attention shift are introduced to the relative position encoding, and merged in a self-attention manner to form a spatial-temporal process similarly to the human visual system. Specifically, we propose a sequence attention shift module to simulate the saccade process, which obtains sequence attention and reduces the computational effort by group attentive convolutional gate recurrent unit. Then, a spatial- temporal correlation encoder module is proposed to encode temporal attention with spatial visual features and obtain the sequential visual features of saccade. Finally, a self-attention fusion module is used to extract the correlation hidden in the relative encoding features. Our proposed AVPT achieves excellent performance on visual sentiment distribution prediction and is comparable to state-of-the-art methods, as demonstrated by extensive experiments on the Flickr_LDL and Twitter_LDL datasets.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心