详细信息

Cognitive Reasoning State Machine for Few-Shot Visual Question Answering  ( EI收录)  

文献类型:期刊文献

英文题名:Cognitive Reasoning State Machine for Few-Shot Visual Question Answering

作者:Wei, Yifan[1]; Liu, Xiaoqiang[2]; Zhang, Jing[1]

机构:[1] Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, China; [2] SenseTime Group, Shanghai, China

年份:2026

外文期刊名:SSRN

收录:EI(收录号:20260143143)

语种:英文

外文关键词:Cognitive systems - Modal analysis - Question answering

摘要:The theory of psychological cognition divides reasoning process of the visual question answering into different states, including seeing, imagining, drawing and answering, which corresponds to understanding, retrieving, reasoning and answering four procedures respectively. Inspired by this, we proposed a Cognitive Reasoning State Machine (CRSM) for few-shot visual question answering, which comprises four modules corresponding to the above reasoning process and realizes effectively cross-modal reasoning by simulating the cognition processing of humans. Each stage of the cognitive process is realized by a Feature Extraction Module (FEM), a novel Prototype Retrieval Module (PRM), a novel Attention Reasoning Module (ARM), and a novel Answer Examining Module (AEM) respectively. Firstly, we filter the key objects of the image in PRM based on the pre-trained object prototypes. Subsequently, the ARM further reinforces the filtering of redundant information and enables fine-grained interactions and cross-modal reasoning between visual and textual features. Finally, an iterative examination of each prediction in the reasoning process is conducted to enhance answer accuracy in AEM. Extensive experiments show that our method consistently outperforms existing approaches, achieving 9.89%–20.25% accuracy gains over the strong MGGN baseline and even larger improvements over the prior state-of-the-art on Toronto COCO-QA, while also establishing new state-of-the-art results on two other few-shot VQA datasets. The code is available at https://github.com/JingVIPLab/CRSM. ? 2026, The Authors. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心