详细信息
Parallel-fusion LSTM with synchronous semantic and visual information for image captioning ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Parallel-fusion LSTM with synchronous semantic and visual information for image captioning
作者:Zhang, Jing[1];Li, Kangkang[1];Wang, Zhe[1]
机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai, Peoples R China
年份:2021
卷号:75
外文期刊名:JOURNAL OF VISUAL COMMUNICATION AND IMAGE REPRESENTATION
收录:;EI(收录号:20210709934626);WOS:【SCI-EXPANDED(收录号:WOS:000633494600005)】;
基金:This research has been supported by the National Nature Science Foundation of China (Grant 61806078).
语种:英文
外文关键词:Image captioning; Parallel-fusion LSTM; Attention mechanism; Guiding LSTM
摘要:For synchronously combining the dynamic semantic and visual information in the decoder part of image captioning, we propose a novel parallel-fusion LSTM (pLSTM) structure in this paper. Two parallel LSTMs with attributes and visual information of image are fused by the hidden states at every time step, which makes the attributes and visual information complementary or enhanced for generating more accurate captions. According to the different ways of integrating semantic information from attribute LSTM to visual LSTM, we propose two models pLSTM with attention (pLSTM-A) and pLSTM with guiding (pLSTM-G). pLSTM-A can automatically capture the crucial semantic and visual information to generate captions, and pLSTM-G directly adjusts the hidden state of visual LSTM by synchronous semantic information to the critical region. For verifying the effectiveness of our proposed pLSTM, we conduct a series of experiments on MSCOCO and Flickr30K datasets, and the experimental results outperform some state-of-the-art image captioning methods.
参考文献:
正在载入数据...
