详细信息
Saccade inspired Attentive Visual Patch Transformer for image sentiment analysis ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Saccade inspired Attentive Visual Patch Transformer for image sentiment analysis
作者:Zhang, Jing[1];Zhu, Jixiang[1];Sun, Han[1];Zhang, Xinzhou[1];Liu, Jiangpei[1]
机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai, Peoples R China
年份:2025
卷号:174
外文期刊名:APPLIED SOFT COMPUTING
收录:;EI(收录号:20251118054492);WOS:【SCI-EXPANDED(收录号:WOS:001450613600001)】;
基金:This study was supported by the Natural Science Foundation Shanghai, China "Research on image sentiment analysis and expression based on human vision and cognitive psychology" (22ZR1418400) .
语种:英文
外文关键词:Image sentiment analysis; Attentive visual patch transformer; Visual attention shift; Saccade mechanism
摘要:The generation of image-evoked emotion is usually regarded as a transient process in the image sentiment analysis. However, according to the saccade mechanism of the human visual system, the evoked emotion generated during the saccade process changes over time and attention. Based on above analysis, we propose an Attentive Visual Patch Transformer (AVPT), using visual attention sequence to represent the sentiment context of images and predict the possible distribution of sentiment. In AVPT, the spatial structure in the form of patches are reconstructed and reorganized by visual attention shift sequentially. Simultaneously, the temporal characteristics of attention shift are introduced to the relative position encoding, and merged in a self-attention manner to form a spatial-temporal process similarly to the human visual system. Specifically, we propose a sequence attention shift module to simulate the saccade process, which obtains sequence attention and reduces the computational effort by group attentive convolutional gate recurrent unit. Then, a spatial- temporal correlation encoder module is proposed to encode temporal attention with spatial visual features and obtain the sequential visual features of saccade. Finally, a self-attention fusion module is used to extract the correlation hidden in the relative encoding features. Our proposed AVPT achieves excellent performance on visual sentiment distribution prediction and is comparable to state-of-the-art methods, as demonstrated by extensive experiments on the Flickr_LDL and Twitter_LDL datasets.
参考文献:
正在载入数据...
