详细信息

An automatic approach for constructing a knowledge base of symptoms in Chinese  ( SCI-EXPANDED收录)  

文献类型:期刊文献

英文题名:An automatic approach for constructing a knowledge base of symptoms in Chinese

作者:Ruan, Tong[1];Wang, Mengjie[1];Sun, Jian[1];Wang, Ting[1];Zeng, Lu[1];Yin, Yichao[2];Gao, Ju[2]

机构:[1]East China Univ Sci & Technol, Shanghai, Peoples R China;[2]Shanghai Shuguang Hosp, Shanghai 200025, Peoples R China

年份:2017

卷号:8

外文期刊名:JOURNAL OF BIOMEDICAL SEMANTICS

收录:;WOS:【SCI-EXPANDED(收录号:WOS:000411384500008)】;

基金:This work and the publication cost of this paper was supported by the 863 plan of China Ministry of Science and Technology (project No: 2015AA020107), "Action Plan for Innovation on Science and Technology" Projects of Shanghai(project No: 16511101000), Research on the Construction Technology of the Healthy and Aged Knowledge Base Based on the Combination of Medical and Care (project No: 2015BAH12F01- 05), and Research on Efficient Query Algorithm for Large Scale Annotated Semantic Knowledge (project No: 61402173).

语种:英文

外文关键词:Knowledge base; Symptoms in Chinese; Linked data; Information extraction

摘要:Background: While a large number of well-known knowledge bases (KBs) in life science have been published as Linked Open Data, there are few KBs in Chinese. However, KBs in Chinese are necessary when we want to automatically process and analyze electronic medical records (EMRs) in Chinese. Of all, the symptom KB in Chinese is the most seriously in need, since symptoms are the starting point of clinical diagnosis. Results: We publish a public KB of symptoms in Chinese, including symptoms, departments, diseases, medicines, and examinations as well as relations between symptoms and the above related entities. To the best of our knowledge, there is no such KB focusing on symptoms in Chinese, and the KB is an important supplement to existing medical resources. Our KB is constructed by fusing data automatically extracted from eight mainstream healthcare websites, three Chinese encyclopedia sites, and symptoms extracted from a larger number of EMRs as supplements. Methods: Firstly, we design data schema manually by reference to the Unified Medical Language System (UMLS). Secondly, we extract entities from eight mainstream healthcare websites, which are fed as seeds to train a multi-class classifier and classify entities from encyclopedia sites and train a Conditional Random Field (CRF) model to extract symptoms from EMRs. Thirdly, we fuse data to solve the large-scale duplication between different data sources according to entity type alignment, entity mapping, and attribute mapping. Finally, we link our KB to UMLS to investigate similarities and differences between symptoms in Chinese and English. Conclusions: As a result, the KB has more than 26,000 distinct symptoms in Chinese including 3968 symptoms in traditional Chinese medicine and 1029 synonym pairs for symptoms. The KB also includes concepts such as diseases and medicines as well as relations between symptoms and the above related entities. We also link our KB to the Unified Medical Language System and analyze the differences between symptoms in the two KBs.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心