详细信息
MKGF: A multi-modal knowledge graph based RAG framework to enhance LVLMs for Medical visual question answering ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:MKGF: A multi-modal knowledge graph based RAG framework to enhance LVLMs for Medical visual question answering
作者:Wu, Yinan[1];Lu, Yuming[1];Zhou, Yan[1];Ding, Yifan[2];Liu, Jingping[1];Ruan, Tong[1]
机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Fudan Univ, Zhongshan Hosp, Dept Crit Care Med, Shanghai 200032, Peoples R China
年份:2025
卷号:635
外文期刊名:NEUROCOMPUTING
收录:;EI(收录号:20251218075281);WOS:【SCI-EXPANDED(收录号:WOS:001452370700001)】;
基金:This work was supported by the National Key Research and Devel-opment Program of China (2023YFF1204904) .
语种:英文
外文关键词:Multi-modal; Knowledge graph; Large language model
摘要:Medical visual question answering (MedVQA) is a challenging task that requires models to understand medical images and return accurate responses for the given questions. Most recent methods focus on transferring general-domain large vision-language models (LVLMs) to the medical domain by constructing medical instruction datasets and in-context learning. However, the performance of these methods are limited due to the hallucination issue of LVLMs. In addition, fine-tuning the abundant parameters of LVLMs on medical instruction datasets is high time and economic cost. Hence, we propose a MKGF framework that leverages a multi-modal medical knowledge graph (MMKG) to relieve the hallucination issue without fine-tuning the abundant parameters of LVLMs. Firstly, we employ a pre-trained text retriever to build question-knowledge relations on training set. Secondly, we train a multi-modal retriever with these relations. Finally, we use it to retrieve question-relevant knowledge and enhance the performance of LVLMs on the test set. To evaluate the effectiveness of MKGF, we conduct extensive experiments on two public datasets Slake and VQA-RAD. Our method improves the pre-trained SOTA LVLMs by 10.15% and 9.32%, respectively. The source codes are available at https://github.com/ehnal/MKGF.
参考文献:
正在载入数据...
