详细信息
PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition ( EI收录)
文献类型:期刊文献
英文题名:PRIMED: Adaptive Modality Suppression for Referring Audio-Visual Segmentation via Biased Competition
作者:He, Yuchen[1]; Zhang, Jing[1]
机构:[1] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China
年份:2026
外文期刊名:arXiv
收录:EI(收录号:20260284618)
语种:英文
外文关键词:Audio acoustics - Benchmarking - Contrastive Learning - Image segmentation - Modal analysis - Visual languages
摘要:Referring Audio-Visual Segmentation (Ref-AVS) seeks to localize and segment target objects in video frames based on visual, auditory, and textual referring cues. The task is challenging because the relevance of different modalities varies across referring expressions and scenes, while existing methods typically treat multimodal cues as homogeneous inputs for fusion, prompting, or reasoning, making them vulnerable to irrelevant or misleading modalities. To address this problem, we propose PRIMED, inspired by the biased competition theory in cognitive neuroscience, which explicitly models both visual perception and language-driven prior modulation, and enables more accurate Ref-AVS by adaptive modality suppression. Specifically, a Modality Prior Decoder first estimates whether the referring expression relies primarily on audio, vision, or their joint interaction, generating a modality prior to adaptively guide high-level attention. A Token Distiller further extracts compact global visual tokens from high-level features and shares them across Competition-aware Cross-modal Fusion modules to provide hierarchical global context. Additionally, we introduce a Spatial-Aware Semantic Alignment loss to further enhance foreground-background discrimination through contrastive learning. Extensive experiments on the Ref-AVS benchmark demonstrate that PRIMED achieves state-of-the-art overall performance. ? 2026, CC BY.
参考文献:
正在载入数据...
