详细信息

RRNorm: A Novel Framework for Chinese Disease Diagnoses Normalization via LLM-Driven Terminology Component Recognition and Reconstruction  ( EI收录)  

文献类型:期刊文献

英文题名:RRNorm: A Novel Framework for Chinese Disease Diagnoses Normalization via LLM-Driven Terminology Component Recognition and Reconstruction

作者:Fan, Yongqi[1]; Zhu, Yansha[1]; Xue, Kui[2]; Liu, Jingping[1]; Ruan, Tong[1]

机构:[1] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China; [2] Shanghai Artificial Intelligence Laboratory, Shanghai, China

年份:2024

起止页码:9162

外文期刊名:Proceedings of the Annual Meeting of the Association for Computational Linguistics

收录:EI(收录号:20244017142407)

语种:英文

外文关键词:Contrastive Learning - Diagnosis - Diseases - Terminology

摘要:The Clinical Terminology Normalization aims at finding standard terms from a given termbase for mentions extracted from clinical texts. However, we found that extracted mentions suffer from the multi-implication problem, especially disease diagnoses. The reason for this is that physicians often use abbreviations, conjunctions, and juxtapositions when writing diagnoses, and it is difficult to manually decompose. To address this problem, we propose a Terminology Component Recognition and Reconstruction strategy that leverages the reasoning capability of large language models (LLMs) to recognize the components of terms, enabling automated decomposition and transforming original mentions into multiple atomic mentions. Furthermore, we adopt the mainstream "Recall and Rank" framework to apply the benefits of the above strategy to the task flow. By leveraging the LLM incorporating the advanced sampling strategies, we design a sampling algorithm for atomic mentions and train the recall model using contrastive learning. Besides the information about the components is also used as knowledge to guide the final term ranking and selection. The experimental results show that our proposed strategy effectively improves the performance of the terminology normalization task and our proposed approach achieves state-of-the-art on the experimental dataset. We release our code and data on the repository RRNorm. ? 2024 Association for Computational Linguistics.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心