详细信息
From electronic health records to terminology base: A novel knowledge base enrichment approach ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:From electronic health records to terminology base: A novel knowledge base enrichment approach
作者:Zhang, Jiaying[1];Zhang, Zhixing[1];Zhang, Huanhuan[1];Ma, Zhiyuan[1];Ye, Qi[1];He, Ping[2];Zhou, Yangming[1]
机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Shanghai Hosp Dev Ctr, Shanghai 200041, Peoples R China
年份:2021
卷号:113
外文期刊名:JOURNAL OF BIOMEDICAL INFORMATICS
收录:;EI(收录号:20204909567446);WOS:【SCI-EXPANDED(收录号:WOS:000615920400001)】;
基金:The authors would like to thank Jun Wang (Shanghai SimMed Technology Limited Company) for providing valuable advices in the construction of clinical indicator terminology base. This work is supported by the National Natural Science Foundation of China (No. 61772201), and National Key Research and Development Program of China for Precision Medical Research (No. 2018YFC0910500).
语种:英文
外文关键词:Knowledge base; Terminology enriching; Entity alignment; Pre-trained language model; Graph convolutional network
摘要:Enriching terminology base (TB) is an important and continuous process, since formal term can be renamed and new term alias emerges all the time. As a potential supplementary for TB enrichment, electronic health record (EHR) is a fundamental source for clinical research and practise. The task to align the set of external terms in EHRs to TB can be regarded as entity alignment without structure information. Conventional approaches mainly use internal structural information of multiple knowledge bases (KBs) to map entities and their counterparts among KBs. However, the external terms in EHRs are independent clinical terms, which lack of interrelations. To achieve entity alignment in this case, we proposed a novel automatic TB enrichment approach, named semantic & structure embeddings-based relevancy prediction (S2ERP). To obtain the semantic embedding of external terms, we fed them with formal entity into a pre-trained language model. Meanwhile, a graph convolutional network was used to obtain the structure embeddings of the synonyms and hyponyms in TB. Afterwards, S2ERP combines both embeddings to measure the relevancy. Experimental results on clinical indicator TB, collected from 38 top-class hospitals of Shanghai Hospital Development Center, showed that the proposed approach outperforms baseline methods by 14.16% in Hits@1.
参考文献:
正在载入数据...
