详细信息

MSDIAGNOSIS: A Benchmark for Evaluating Large Language Models in Multi-Step Clinical Diagnosis  ( EI收录)  

文献类型:期刊文献

英文题名:MSDIAGNOSIS: A Benchmark for Evaluating Large Language Models in Multi-Step Clinical Diagnosis

作者:Hou, Ruihui[1]; Chen, Shencheng[1]; Fan, Yongqi[1]; Yu, Guangya[1]; Zhu, Lifeng[2]; Sun, Jing[2]; Liu, Jingping[1]; Ruan, Tong[1]

机构:[1] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China; [2] Ruijin Hospital, Shanghai Jiao Tong University, School of Medicine, Shanghai, China

年份:2024

外文期刊名:arXiv

收录:EI(收录号:20240356265)

语种:英文

外文关键词:Benchmarking

摘要:Clinical diagnosis is critical in medical practice, typically requiring a continuous and evolving process that includes primary diagnosis, differential diagnosis, and final diagnosis. However, most existing clinical diagnostic tasks are single-step processes, which does not align with the complex multi-step diagnostic procedures found in real-world clinical settings. In this paper, we propose a Chinese clinical diagnostic benchmark, called MSDiagnosis. This benchmark consists of 2,225 cases from 12 departments, covering tasks such as primary diagnosis, differential diagnosis, and final diagnosis. Additionally, we propose a novel and effective framework. This framework combines forward inference, backward inference, reflection, and refinement, enabling the large language model to self-evaluate and adjust its diagnostic results. To this end, we test open-source models, closed-source models, and our proposed framework. The experimental results demonstrate the effectiveness of the proposed method. We also provide a comprehensive experimental analysis and suggest future research directions for this task. Copyright ? 2024, The Authors. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心