详细信息
MedEureka: A Medical Domain Benchmark for Multi-Granularity and Multi-Data-Type Embedding-Based Retrieval ( EI收录)
文献类型:期刊文献
英文题名:MedEureka: A Medical Domain Benchmark for Multi-Granularity and Multi-Data-Type Embedding-Based Retrieval
作者:Fan, Yongqi[1]; Wang, Nan[1]; Xue, Kui[2]; Liu, Jingping[1]; Ruan, Tong[1]
机构:[1] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China; [2] Intelligent Healthcare, Shanghai Artificial Intelligence Laboratory, Shanghai, China
年份:2025
起止页码:2825
外文期刊名:2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Proceedings of the Conference Findings, NAACL 2025
收录:EI(收录号:20260519997630)
语种:英文
外文关键词:Benchmarking - Computational linguistics - Electronic health record - Information retrieval - Medical computing - Natural language processing systems - Statistical tests
摘要:Embedding-based retrieval (EBR), the mainstream approach in information retrieval (IR), aims to help users obtain relevant information and plays a crucial role in retrieval-augmented generation (RAG) techniques of large language models (LLMs). Numerous methods have been proposed to significantly improve the quality of retrieved content and many generic benchmarks are proposed to evaluate the retrieval abilities of embedding models. However, texts in the medical domain present unique contexts, structures, and language patterns, such as terminology, doctor-patient dialogue, and electronic health records (EHRs). Despite these unique features, specific benchmarks for medical context retrieval are still lacking. In this paper, we propose MedEureka, an enriched benchmark designed to evaluate medical-context retrieval capabilities of embedding models with multi-granularity and multi-data types. MedEureka includes four levels of granularity and six types of medical texts, encompassing 18 datasets, incorporating granularity and data type description to prompt instruction-fine-tuned text embedding models for embedding generation. We also provide the MedEureka Toolkit to support evaluation on the MedEureka test set. Our experiments evaluate state-of-the-art open-source and proprietary embedding models, and fine-tuned classical baselines, providing a detailed performance analysis. This underscores the challenges of using embedding models for medical domain retrieval and the need for further research. Our code and data are released in the repository: https://github com/JOHNNY-fans/MedEureka. ? 2025 Association for Computational Linguistics.
参考文献:
正在载入数据...
