详细信息
Multi-feature fusion enhanced transformer with multi-layer fused decoding for image captioning ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:Multi-feature fusion enhanced transformer with multi-layer fused decoding for image captioning
作者:Zhang, Jing[1];Fang, Zhongjun[1];Wang, Zhe[1]
机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai, Peoples R China
年份:2023
卷号:53
期号:11
起止页码:13398
外文期刊名:APPLIED INTELLIGENCE
收录:;EI(收录号:20224212901323);WOS:【SCI-EXPANDED(收录号:WOS:000865724500001)】;
基金:This study was funded by the Nature Science Foundation of Shanghai "Research on image sentiment analysis and expression based on human vision and cognitive psychology"(grant number 22ZR1418400).
语种:英文
外文关键词:Image captioning; Multi-feature fusion enhanced transformer; Multi-layer fused decoding
摘要:The objects' semantic information of the image is vital for image captioning. Though some methods have used semantic information, the alignment between the specific semantic feature and the corresponding visual feature has not been explored. In this paper, we propose a novel Multi-Feature Fusion enhanced Transformer (MFF-Transformer) for image captioning, which can realize the multi-feature fusion by aligning the specific semantic features with their corresponding visual features for achieving a better visual feature representation. In the MFF-Transformer, a novel Interacted Multi-Feature REpresentation (IMFRE) module is proposed, which effectively fuses the visual features with the semantic features by adaptive pooling to obtain global and local multi-feature information for enhancing the visual features. In addition, a new Multi-Layer Features Fusion (MLFF) module is proposed to achieve complete and valid multi-layer decoding information for utilizing the hierarchical context of the MFF-Transformer model. Experiments on the MSCOCO dataset illustrate that our proposed MFF-Transformer can achieve good performance and outperform other state-of-the-art methods.
参考文献:
正在载入数据...
