详细信息

MKGF: A multi-modal knowledge graph based RAG framework to enhance LVLMs for Medical visual question answering  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:MKGF: A multi-modal knowledge graph based RAG framework to enhance LVLMs for Medical visual question answering

作者:Wu, Yinan[1];Lu, Yuming[1];Zhou, Yan[1];Ding, Yifan[2];Liu, Jingping[1];Ruan, Tong[1]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Fudan Univ, Zhongshan Hosp, Dept Crit Care Med, Shanghai 200032, Peoples R China

年份:2025

卷号:635

外文期刊名:NEUROCOMPUTING

收录:;EI(收录号:20251218075281);WOS:【SCI-EXPANDED(收录号:WOS:001452370700001)】;

基金:This work was supported by the National Key Research and Devel-opment Program of China (2023YFF1204904) .

语种:英文

外文关键词:Multi-modal; Knowledge graph; Large language model

摘要:Medical visual question answering (MedVQA) is a challenging task that requires models to understand medical images and return accurate responses for the given questions. Most recent methods focus on transferring general-domain large vision-language models (LVLMs) to the medical domain by constructing medical instruction datasets and in-context learning. However, the performance of these methods are limited due to the hallucination issue of LVLMs. In addition, fine-tuning the abundant parameters of LVLMs on medical instruction datasets is high time and economic cost. Hence, we propose a MKGF framework that leverages a multi-modal medical knowledge graph (MMKG) to relieve the hallucination issue without fine-tuning the abundant parameters of LVLMs. Firstly, we employ a pre-trained text retriever to build question-knowledge relations on training set. Secondly, we train a multi-modal retriever with these relations. Finally, we use it to retrieve question-relevant knowledge and enhance the performance of LVLMs on the test set. To evaluate the effectiveness of MKGF, we conduct extensive experiments on two public datasets Slake and VQA-RAD. Our method improves the pre-trained SOTA LVLMs by 10.15% and 9.32%, respectively. The source codes are available at https://github.com/ehnal/MKGF.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心