详细信息
Adaptive Semantic-Enhanced Transformer for Image Captioning ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Adaptive Semantic-Enhanced Transformer for Image Captioning
作者:Zhang, Jing[1];Fang, Zhongjun[1];Sun, Han[1];Wang, Zhe[1]
机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Technol, Shanghai 200237, Peoples R China
年份:2024
卷号:35
期号:2
起止页码:1785
外文期刊名:IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS
收录:;EI(收录号:20222812350123);WOS:【SCI-EXPANDED(收录号:WOS:000824712700001)】;
基金:This work was supported in part by the Shanghai Science and Technology Program-Distributed and generative few-shot algorithm and theory research-under Grant 20511100600 and in part by the Natural Science Foundation of Shanghai-Research on image sentiment analysis and expression based on human vision and cognitive psychology-under Grant 22ZR1418400.
语种:英文
外文关键词:Semantics; Visualization; Decoding; Transformers; Adaptation models; Detectors; Image reconstruction; Adaptive gated mechanism (AGM); adaptive semantic-enhanced transformer (AS-Transformer); constrained weakly supervised learning; image captioning
摘要:In the research on image captioning, rich semantic information is very important for generating critical caption words as guiding information. However, semantic information from offline object detectors involves many semantic objects that do not appear in the caption, thereby bringing noise into the decoding process. To produce more accurate semantic guiding information and further optimize the decoding process, we propose an end-to-end adaptive semantic-enhanced transformer (AS-Transformer) model for image captioning. For semantic enhancement information extraction, we propose a constrained weaklysupervised learning (CWSL) module, which reconstructs the semantic object's probability distribution detected by the multiple instances learning (MIL) through a joint loss function. These strengthened semantic objects from the reconstructed probability distribution can better depict the semantic meaning of images. Also, for semantic enhancement decoding, we propose an adaptive gated mechanism (AGM) module to adjust the attention between visual and semantic information adaptively for the more accurate generation of caption words. Through the joint control of the CWSL module and AGM module, our proposed model constructs a complete adaptive enhancement mechanism from encoding to decoding and obtains visual context that is more suitable for captions. Experiments on the public Microsoft Common Objects in COntext (MSCOCO) and Flickr30K datasets illustrate that our proposed AS-Transformer can adaptively obtain effective semantic information and adjust the attention weights between semantic and visual information automatically, which achieves more accurate captions compared with semantic enhancement methods and outperforms state-of-the-art methods.
参考文献:
正在载入数据...
