详细信息

Multimodal Coupling Prompt Learning for Image Classification Tasks  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Multimodal Coupling Prompt Learning for Image Classification Tasks

作者:Liu, Yufei[1];Cheng, Hua[1];Fang, Yiquan[1];Pan, Yiming[1];Qian, Zehong[1];Chen, Xiaoning[1]

机构:[1]East China Univ Sci & Technol, Dept Informat Sci & Engn, Shanghai, Peoples R China

年份:2025

卷号:38

期号:4

起止页码:684

外文期刊名:EUROPEAN JOURNAL ON ARTIFICIAL INTELLIGENCE

收录:;EI(收录号:20253919248124);WOS:【SCI-EXPANDED(收录号:WOS:001492609800001)】;

语种:英文

外文关键词:prompt learning; image classification; cross-attention; CLIP

摘要:In recent years, vision-language pretraining (VLP) models have become a crucial driving force in the advancement of artificial intelligence. Besides, studies such as contrastive language-image pretraining (CLIP) have demonstrated that incorporating prompt learning within VLP models can significantly enhance the performance of downstream tasks. However, we believe that CLIP's visual encoder suffers from feature extraction bias in image classification tasks, which is because of the uneven quantity and distribution of image features CLIP learned between the pretraining and fine-tuning stages. This can be further summarized as an inherent bias in feature extraction for differently distributed samples during the pretraining phase. To address the above problem this paper proposes (i) text-semantic hierarchical injection prompt learning method, which constructs self-attention layers and prompt mapping structures and injects text semantic features into the visual encoder layer by layer to generate visual prompt features and (ii) visual-semantic attention interactive prompt learning method, which further integrates text embeddings with the output features of the visual encoder through cross-attention and constructs instance-level text prompt features for each image. Based on the two above methods, this paper further proposes the multimodal coupling prompt learning CLIP (MCPL-CLIP) to enhance CLIP's performance in image classification tasks. Experiments conducted on 15 image classification datasets demonstrate that MCPL-CLIP outperforms baseline models such as MaPLe, CoCoOp, and CoOp in cross-dataset transfer, domain generalization, and base-to-novel class generalization tasks, showcasing its superior text semantic representation and visual feature extraction capabilities.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心