详细信息

Enhancing Visual In-Context Learning by Multi-Faceted Fusion  ( EI收录)  

文献类型:期刊文献

英文题名:Enhancing Visual In-Context Learning by Multi-Faceted Fusion

作者:Liao, Wenwen[1]; Yu, Jianbo[2]; Wang, Yuansong[3]; Jiang, Qingchao[4]; Yang, Xiaofeng[2]

机构:[1] College of Intelligent Robotics and Advance Manufacturing, Fudan University, China; [2] School of Microelectronics, Fudan University, China; [3] Tsinghua Shenzhen International Graduate School, Tsinghua University, China; [4] School of Information Science and Engineering, East China University of Science and Technology, China

年份:2026

外文期刊名:arXiv

收录:EI(收录号:20260092029)

语种:英文

外文关键词:Collaborative learning - Fusion reactions - Image fusion - Object detection

摘要:Visual In-Context Learning (VICL) has emerged as a powerful paradigm, enabling models to perform novel visual tasks by learning from in-context examples. The dominant "retrieve-then-prompt" approach typically relies on selecting the single best visual prompt, a practice that often discards valuable contextual information from other suitable candidates. While recent work has explored fusing the top-K prompts into a single, enhanced representation, this still simply collapses multiple rich signals into one, limiting the model’s reasoning capability. We argue that a more multi-faceted, collaborative fusion is required to unlock the full potential of these diverse contexts. To address this limitation, we introduce a novel framework that moves beyond single-prompt fusion towards an multi-combination collaborative fusion. Instead of collapsing multiple prompts into one, our method generates three contextual representation branches, each formed by integrating information from different combinations of top-quality prompts. These complementary guidance signals are then fed into proposed MULTI-VQGAN architecture, which is designed to jointly interpret and utilize collaborative information from multiple sources. Extensive experiments on diverse tasks, including foreground segmentation, single-object detection, and image colorization, highlight its strong cross-task generalization, effective contextual fusion, and ability to produce more robust and accurate predictions than existing methods. Code will be released after acceptance. Copyright ? 2026, The Authors. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心