详细信息
Visual enhanced gLSTM for image captioning ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Visual enhanced gLSTM for image captioning
作者:Zhang, Jing[1];Li, Kangkang[1];Wang, Zhenkun[1];Zhao, Xianwen[1];Wang, Zhe[1]
机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China
年份:2021
卷号:184
外文期刊名:EXPERT SYSTEMS WITH APPLICATIONS
收录:;EI(收录号:20212810636302);WOS:【SCI-EXPANDED(收录号:WOS:000708093400015)】;
基金:This research has been supported by the National Nature Science Foundation of China (Grant 61806078 and 61672227).
语种:英文
外文关键词:Image caption; Visual enhanced-gLSTM; Bag of; Region of interest; Salient region
摘要:For reducing the negative impact of the gradient diminishing on guiding long-short term memory (gLSTM) model in image captioning, we propose a visual enhanced gLSTM model for image caption generation. In this paper, the visual features of image's region of interest (RoI) are extracted and used as guiding information in gLSTM, in which visual information of RoI is added to gLSTM for generating more accurate image captions. Two visual enhanced methods based on region and entire image are proposed respectively. Among them the visual features from the important semantic region by CNN and the full image visual features by visual words are extracted to guide the LSTM for generating the most important semantic words. Then the visual features and text features of similar images are respectively projected to the common semantic space to obtain visual enhancement guiding information by canonical correlation analysis, and added to each memory cell of gLSTM for generating caption words. Compared with the original gLSTM method, visual enhanced gLSTM model focuses on important semantic region, which is more in line with human perception of images. Experiments on Flickr8k dataset illustrate that the proposed method can achieve more accurate image captions, and outperform the baseline gLSTM algorithm and other popular image captioning methods.
参考文献:
正在载入数据...
