详细信息

Improving Image Captioning through Visual and Semantic Mutual Promotion  ( EI收录)  

文献类型:期刊文献

英文题名:Improving Image Captioning through Visual and Semantic Mutual Promotion

作者:Zhang, Jing[1]; Xie, Yingshuai[1]; Liu, Xiaoqiang[1]

机构:[1] East China University of Science and Technology, China

年份:2023

起止页码:4716

外文期刊名:MM 2023 - Proceedings of the 31st ACM International Conference on Multimedia

收录:EI(收录号:20235015224405)

语种:英文

外文关键词:Decoding - Deep learning - Image enhancement - Image fusion - Object detection - Statistical tests

摘要:Current image captioning methods commonly use semantic attributes extracted by an object detector to guide visual representation, leaving the mutual guidance and enhancement between vision and semantics under-explored. Neurological studies have revealed that the visual cortex of the brain plays a crucial role in recognizing visual objects, while the prefrontal cortex is involved in the integration of contextual semantics. Inspired by the above studies, we propose a novel Visual-Semantic Transformer (VST) to model the neural interaction between vision and semantics, which explores the mechanism of deep fusion and mutual promotion of multimodal information, realizing more accurate image captioning. To better facilitate the complementary strengths between visual objects and semantic contexts, we propose a global position-sensitive co-attention encoder to realize globally associative, position-aware visual and semantic co-interaction through a mutual cross-attention mechanism. In addition, a multimodal mixed attention module is proposed in the decoder, which achieves adaptive multimodal feature fusion for enhancing the decoding capability. Experimental evidence shows that our VST significantly surpasses the state-of-the-art approaches on MSCOCO dataset and reaches the excellent CIDEr score of 142% on the Karpathy test split. ? 2023 ACM.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心