详细信息
基于空间和多层级联合编码的图像描述算法
Spatial Encoding and Multi-layer Joint Encoding Enhanced Transformer for Image Captioning
文献类型:期刊文献
中文题名:基于空间和多层级联合编码的图像描述算法
英文题名:Spatial Encoding and Multi-layer Joint Encoding Enhanced Transformer for Image Captioning
作者:方仲俊[1,2];张静[1];李冬冬[1,2]
机构:[1]华东理工大学信息科学与工程学院,上海200237;[2]苏州大学江苏省计算机信息处理技术重点实验室,江苏苏州215031
年份:2022
卷号:49
期号:10
起止页码:151
中文期刊名:计算机科学
外文期刊名:Computer Science
收录:CSTPCD;;北大核心:【北大核心2020】;CSCD:【CSCD_E2021_2022】;
基金:国家自然科学基金(61806078)。
语种:中文
中文关键词:图像描述;Transformer;空间编码机制;多层级联合编码机制;注意力机制
外文关键词:Image captioning;Transformer;Spatial encoding mechanism;Multi-level joint encoding mechanism;Attention mechanism
摘要:图像描述是图像理解领域的热点研究课题之一,它是结合计算机视觉和自然语言处理的跨媒体数据分析任务,通过理解图像内容并生成语义和语法都正确的句子来描述图像。现有的图像描述方法多采用编码器-解码器模型,该类方法在提取图像中的视觉对象特征时大多忽略了视觉对象之间的相对位置关系,但它对于正确描述图像的内容是非常重要的。基于此,提出了基于Transformer的空间和多层级联合编码的图像描述方法。为了更好地利用图像中所包含的对象的位置信息,提出了视觉对象的空间编码机制,将各个视觉对象独立的空间关系转换为视觉对象间的相对空间关系,以此来帮助模型识别各个视觉对象间的相对位置关系。同时,在视觉对象的编码阶段,顶部的编码特征保留了更多的贴合图像语义信息,但丢失了图像部分视觉信息,考虑到这一点,文中提出了多层级联合编码机制,通过整合各个浅层的编码层所包含的图像特征信息来完善顶部编码层所蕴含的语义的信息,从而获取到更丰富的贴合图像的语义信息的编码特征。文中在MSCOCO数据集上使用多种评估指标(BLEU,METEOR,ROUGE-L和CIDEr等)对提出的图像描述方法进行评估,并通过消融实验证明了提出的基于空间的编码机制以及多层级联合编码机制能够辅助产生更为准确有效的图像描述语句。对比实验结果表明,所提方法能够产生准确、有效的图像描述并优于大多数最新的算法。
Image captioning is one of the hot research topics in the field of computer vision.It is a cross-media data analysis task that combines computer vision and natural language processing.It describes the image by understanding the content of the image and generating captions that are both semantically and grammatically correct.Existing image captioning methods mostly use the encoder-decoder model.This kind of methods mostly ignore the relative position relationship between visual objects when extracting the visual object features in image, and the relative position relationship between objects is very important for generating accurate captioning.Based on this, this paper proposes a spatial encoding and multi-layer joint encoding enhanced transformer for image captioning.In order to make better use of the position information contained in the image, this paper proposes a spatial encoding mechanism for visual objects, which converts the independent spatial relationship of each visual object into the relative spatial relationship between visual objects to help the model to recognize the relative spatial relationship between each visual object.At the same time, in the encoder part of visual objects, the top encoding feature retains more semantic information that fits the image but loses part of the visual information of the image.Taking this into account, this paper proposes a multi-level joint encoding mechanism to improve the semantic information contained in the top encoding layer by integrating the image feature information contained in each shallow encoding layer, so as to obtain richer semantic features that fit the image.This paper evaluates the proposed image captioning method by multiple evaluation indicators(BLEU,METEOR,ROUGE-L,CIDEr, etc.) on the MSCOCO dataset.The ablation experiment proves that the spatial encoding mechanism and the multi-level joint encoding mechanism proposed in this paper can be helpful in generating more accurate and effective image captions.Comparative experimental results show that the proposed method in can produce accurate and effective image caption and is superior to most of the latest methods.
参考文献:
正在载入数据...
