详细信息

MedOdyssey: A Medical Domain Benchmark for Long Context Evaluation Up to 200K Tokens  ( EI收录)  

文献类型:期刊文献

英文题名:MedOdyssey: A Medical Domain Benchmark for Long Context Evaluation Up to 200K Tokens

作者:Fan, Yongqi[1]; Sun, Hongli[1]; Xue, Kui[2]; Zhang, Xiaofan[3]; Zhang, Shaoting[2]; Ruan, Tong[1]

机构:[1] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China; [2] Intelligent Healthcare, Shanghai Artificial Intelligence Laboratory, Shanghai, China; [3] School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, Shanghai, China

年份:2024

外文期刊名:arXiv

收录:EI(收录号:20240277480)

语种:英文

外文关键词:Computational linguistics

摘要:Numerous advanced Large Language Models (LLMs) now support context lengths up to 128K, and some extend to 200K. Some benchmarks in the generic domain have also followed up on evaluating long-context capabilities. In the medical domain, tasks are distinctive due to the unique contexts and need for domain expertise, necessitating further evaluation. However, despite the frequent presence of long texts in medical scenarios, evaluation benchmarks of long-context capabilities for LLMs in this field are still rare. In this paper, we propose MedOdyssey, the first medical long-context benchmark with seven length levels ranging from 4K to 200K tokens. MedOdyssey consists of two primary components: the medical-context "needles in a haystack" task and a series of tasks specific to medical applications, together comprising 10 datasets. The first component includes challenges such as counter-intuitive reasoning and novel (unknown) facts injection to mitigate knowledge leakage and data contamination of LLMs. The second component confronts the challenge of requiring professional medical expertise. Especially, we design the "Maximum Identical Context" principle to improve fairness by guaranteeing that different LLMs observe as many identical contexts as possible. Our experiment evaluates advanced proprietary and open-source LLMs tailored for processing long contexts and presents detailed performance analyses. This highlights that LLMs still face challenges and need for further research in this area. Our code and data are released in the repository: https://github.com/JOHNNY-fans/MedOdyssey. ? 2024, CC BY.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心