详细信息

Improving Image Captioning through Visual and Semantic Mutual Promotion  ( CPCI-S收录)  

文献类型:会议论文

英文题名:Improving Image Captioning through Visual and Semantic Mutual Promotion

作者:Zhang, Jing[1];Xie, Yingshuai[1];Liu, Xiaoqiang[1]

机构:[1]East China Univ Sci & Technol, Shanghai, Peoples R China

会议论文集:31st ACM International Conference on Multimedia (MM)

会议日期:OCT 29-NOV 03, 2023

会议地点:Ottawa, CANADA

语种:英文

外文关键词:Image Captioning; Transformer; Co-attention; Multimodal Fusion

摘要:Current image captioning methods commonly use semantic attributes extracted by an object detector to guide visual representation, leaving the mutual guidance and enhancement between vision and semantics under-explored. Neurological studies have revealed that the visual cortex of the brain plays a crucial role in recognizing visual objects, while the prefrontal cortex is involved in the integration of contextual semantics. Inspired by the above studies, we propose a novel Visual-Semantic Transformer (VST) to model the neural interaction between vision and semantics, which explores the mechanism of deep fusion and mutual promotion of multimodal information, realizing more accurate image captioning. To better facilitate the complementary strengths between visual objects and semantic contexts, we propose a global position-sensitive co-attention encoder to realize globally associative, position-aware visual and semantic co-interaction through a mutual cross-attention mechanism. In addition, a multimodal mixed attention module is proposed in the decoder, which achieves adaptive multimodal feature fusion for enhancing the decoding capability. Experimental evidence shows that our VST significantly surpasses the state-of-the-art approaches on MSCOCO dataset and reaches the excellent CIDEr score of 142% on the Karpathy test split.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心