详细信息
MSDIAGNOSIS: A Benchmark for Evaluating Large Language Models in Multi-Step Clinical Diagnosis ( EI收录)
文献类型:期刊文献
英文题名:MSDIAGNOSIS: A Benchmark for Evaluating Large Language Models in Multi-Step Clinical Diagnosis
作者:Hou, Ruihui[1]; Chen, Shencheng[1]; Fan, Yongqi[1]; Yu, Guangya[1]; Zhu, Lifeng[2]; Sun, Jing[2]; Liu, Jingping[1]; Ruan, Tong[1]
机构:[1] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China; [2] Ruijin Hospital, Shanghai Jiao Tong University, School of Medicine, Shanghai, China
年份:2024
外文期刊名:arXiv
收录:EI(收录号:20240356265)
语种:英文
外文关键词:Benchmarking
摘要:Clinical diagnosis is critical in medical practice, typically requiring a continuous and evolving process that includes primary diagnosis, differential diagnosis, and final diagnosis. However, most existing clinical diagnostic tasks are single-step processes, which does not align with the complex multi-step diagnostic procedures found in real-world clinical settings. In this paper, we propose a Chinese clinical diagnostic benchmark, called MSDiagnosis. This benchmark consists of 2,225 cases from 12 departments, covering tasks such as primary diagnosis, differential diagnosis, and final diagnosis. Additionally, we propose a novel and effective framework. This framework combines forward inference, backward inference, reflection, and refinement, enabling the large language model to self-evaluate and adjust its diagnostic results. To this end, we test open-source models, closed-source models, and our proposed framework. The experimental results demonstrate the effectiveness of the proposed method. We also provide a comprehensive experimental analysis and suggest future research directions for this task. Copyright ? 2024, The Authors. All rights reserved.
参考文献:
正在载入数据...
