详细信息

Cross-Modal Feature Distribution Calibration for Few-Shot Visual Question Answering  ( EI收录)  

文献类型:期刊文献

英文题名:Cross-Modal Feature Distribution Calibration for Few-Shot Visual Question Answering

作者:Zhang, Jing[1]; Liu, Xiaoqiang[1]; Chen, Mingzhe[1]; Wang, Zhe[1]

机构:[1] Department of Computer Science and Engineering, East China University of Science and Technology, China

年份:2024

卷号:38

期号:7

起止页码:7151

外文期刊名:Proceedings of the AAAI Conference on Artificial Intelligence

收录:EI(收录号:20241515870407)

语种:英文

外文关键词:Artificial intelligence - Benchmarking - Classification (of information) - Spatial distribution

摘要:Few-shot Visual Question Answering (VQA) realizes few-shot cross-modal learning, which is an emerging and challenging task in computer vision. Currently, most of the few-shot VQA methods are confined to simply extending few-shot classification methods to cross-modal tasks while ignoring the spatial distribution properties of multimodal features and cross-modal information interaction. To address this problem, we propose a novel Cross-modal feature Distribution Calibration Inference Network (CDCIN) in this paper, where a new concept named visual information entropy is proposed to realize multimodal features distribution calibration by cross-modal information interaction for more effective few-shot VQA. Visual information entropy is a statistical variable that represents the spatial distribution of visual features guided by the question, which is aligned before and after the reasoning process to mitigate redundant information and improve multi-modal features by our proposed visual information entropy calibration module. To further enhance the inference ability of cross-modal features, we additionally propose a novel pre-training method, where the reasoning sub-network of CDCIN is pretrained on the base class in a VQA classification paradigm and fine-tuned on the few-shot VQA datasets. Extensive experiments demonstrate that our proposed CDCIN achieves excellent performance on few-shot VQA and outperforms state-of-the-art methods on three widely used benchmark datasets. Copyright ? 2024, Association for the Advancement of Artificial Intelligence (www.aaai.org). All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心