详细信息

Hierarchical decoding with latent context for image captioning  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Hierarchical decoding with latent context for image captioning

作者:Zhang, Jing[1];Xie, Yingshuai[1];Li, Kangkang[1];Wang, Zhe[1];Du, Wen[2]

机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China;[2]DS Informat Technol Co Ltd, Shanghai 200032, Peoples R China

年份:2023

卷号:35

期号:3

起止页码:2429

外文期刊名:NEURAL COMPUTING & APPLICATIONS

收录:;EI(收录号:20223612686139);WOS:【SCI-EXPANDED(收录号:WOS:000846241800002)】;

基金:This work was supported by the Shanghai Science and Technology Program "Distributed and generative few-shot algorithm and theory research" under Grant 20511100600 and the Natural Science Foundation of Shanghai "Research on image sentiment analysis and expression based on human vision and cognitive psychology" under Grant 22ZR1418400.

语种:英文

外文关键词:Image captioning; Visual context; Latent context generation; Hierarchical decoding

摘要:Mining more rich visual features and analyzing the context information from image for decoding part has become a challenging problem in image captioning. Some recent works employ other knowledge bases to obtain the additional objects semantic relationships by constructing scene graph, which spend much time on pre-training scene graph and these artificial defined relationships may not be comprehensive. In this paper, a novel hierarchical decoding with latent context method is proposed for image captioning, which analyzes the visual context information and decodes multi-level visual features by a hierarchical decoding method to achieve more accurate caption words. In our proposed method, a novel Latent Context Generation Network (LCGN) is proposed to infer latent relationships between objects without any external knowledge, and meanwhile, a context vector which contains rich neighbor information for each object is constructed. Then a graph convolutional network with attention is used to further aggregate latent context information for achieving high-level context features by combining objects features and their context vectors. Finally, hierarchical decoding based on Triple Long Short-Term Memory (Tri-LSTM) is proposed to decode global features, local features and object features hierarchically, which gradually analyzes the content of the image from the whole to the local to the object. Experiments on MSCOCO dataset prove that our proposed method can achieve extremely competitive results in image captioning and outperform most CNN-RNN architecture methods.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心