详细信息
文献类型:期刊文献
中文题名:基于语义分割的全卷积图像描述模型
英文题名:Fully convolutional image description model based on semantic segmentation
作者:李永生[1];颜秉勇[1];周家乐[1]
机构:[1]华东理工大学信息科学与工程学院,上海200237
年份:2023
卷号:44
期号:1
起止页码:210
中文期刊名:计算机工程与设计
外文期刊名:Computer Engineering and Design
收录:CSTPCD;;北大核心:【北大核心2020】;
基金:国家自然科学基金青年基金项目(61906068)。
语种:中文
中文关键词:图像描述;语义分割;卷积神经网络;编码器;语义信息;长短时记忆网络;解码速度
外文关键词:image description;semantic segmentation;convolutional neural network;encoder;semantic information;long and short-term memory network;decoding speed
摘要:为快速生成准确描述图片内容的语句,提出语义分割和卷积神经网络(convolutional neural network,CNN)相结合的图像描述方法。将图像分类模型和语义分割模型结合为编码器,增强对图像语义信息的利用,采用CNN代替长短时记忆网络(long short term memory,LSTM)作为解码器生成完整描述性语句。通过在MSCOCO数据集上与5种主流算法的对比实验可知,以CNN作为解码器能够大幅提高解码速度,语义信息的增强能够有效提高实验精度,验证了该方法的有效性和可行性。
To quickly generate sentences that accurately describe the content of a picture,an image description method combining semantic segmentation and convolutional neural network(CNN)was proposed.The image classification model and semantic segmentation model were combined into an encoder to enhance the use of image semantic information,and CNN was used instead of long short term memory(LSTM)as a decoder to generate complete descriptive sentences.By comparing experiments with five mainstream algorithms on the MSCOCO data set,it can be seen that using CNN as a decoder can greatly increase the decoding speed,and the enhancement of semantic information can also effectively improve the experimental accuracy,which verifies the effectiveness and feasibility of the method.
参考文献:
正在载入数据...
