详细信息
Cross-modal heterogeneous graph reasoning network for visual question answering ( EI收录)
文献类型:期刊文献
英文题名:Cross-modal heterogeneous graph reasoning network for visual question answering
作者:Zhang, Jing[1]; Teng, Jiong[1]; Ding, Weichao[1]; Wang, Zhe[1]
机构:[1] Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, China
年份:2025
卷号:37
期号:22
起止页码:17701
外文期刊名:Neural Computing and Applications
收录:EI(收录号:20251918388503)
语种:英文
外文关键词:Modal analysis
摘要:Most current Visual Question Answering (VQA) methods struggle to achieve effective cross-modal interaction between visual and semantic information, resulting in difficulties in accurately combining visual content with contextual semantics for answer prediction. To address this problem, a Cross-modal Heterogeneous Graph Reasoning Network (CHGRN) is proposed for VQA, incorporating a novel Cross-modal Reasoning Module (CRM) to enhance the interactive analysis between images and questions, enabling more profound joint reasoning of visual and semantic features. The CRM improves cross-modal information reasoning by effectively analyzing and understanding the visual semantic information in the target regions of images. Additionally, an answer type prediction module is introduced, employing multi-task learning with answer type annotations to filter out irrelevant semantic information, thereby improving reasoning accuracy. Moreover, the semantic-assisted attention-aligned decoder ensures precise alignment between visual and semantic data. Extensive experiments demonstrate that the proposed CHGRN achieves excellent performance in visual question answering and outperforms most state-of-the-art methods on widely used public datasets. ? The Author(s), under exclusive licence to Springer-Verlag London Ltd., part of Springer Nature 2025.
参考文献:
正在载入数据...
