详细信息

Region-Attention Prompt Learning for CLIP  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Region-Attention Prompt Learning for CLIP

作者:Pan, Yiming[1];Cheng, Hua[1];Fang, Yiquan[1];Liu, Yufei[1]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai, Peoples R China

年份:2023

卷号:45

期号:5

起止页码:7221

外文期刊名:JOURNAL OF INTELLIGENT & FUZZY SYSTEMS

收录:;EI(收录号:20234615072586);WOS:【SCI-EXPANDED(收录号:WOS:001099536200003)】;

语种:英文

外文关键词:Prompt learning; CLIP; Cross-Attention mechanism; image classfication

摘要:Pre-trained Visual Language Models (VLMs) like CLIP have shown great potential in the multimodal domain. Among this, using different modal contexts and interaction features to construct prompt can stimulate the model's prior knowledge circuit more accurately, thus generating better outputs. However, in CLIP, the formal mismatch of textual descriptions between the pre-training and inference phases results in a suboptimal representation ability of prompt, which is detrimental to model alignment learning. Therefore, Region-Attention Prompt (RAP) is proposed, which introduces region features to enrich the semantic representation of prompt. RAP is acquired by the Cross-Attention mechanism between images and texts, and it is essentially a region-level prompt with category-sensitive properties. For each category, RAP adaptively assigns greater attention weight to image regions that are more semantically relevant to the category. Besides, CLIP is equipped with RAP (called RA-CLIP) to improve image classification performance in generalization scenarios. Extensive experiments demonstrate that RA-CLIP outperforms the current SOTA CoCoOp 0.4% - 4.16% on base classes and 0.25% - 11.34% on new classes, across 7 datasets. In addition, we show that focusing on category-related regions to construct prompt can further improve the model's alignment ability.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心