详细信息

NE-LP: Normalized entropy- and loss prediction-based sampling for active learning in Chinese word segmentation on EHRs  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:NE-LP: Normalized entropy- and loss prediction-based sampling for active learning in Chinese word segmentation on EHRs

作者:Cai, Tingting[1];Ma, Zhiyuan[2];Zheng, Hong[1];Zhou, Yangming[1]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai 200237, Peoples R China;[2]Univ Shanghai Sci & Technol, Inst Machine Intelligence, Shanghai 200093, Peoples R China

年份:2021

卷号:33

期号:19

起止页码:12535

外文期刊名:NEURAL COMPUTING & APPLICATIONS

收录:;EI(收录号:20211410162331);WOS:【SCI-EXPANDED(收录号:WOS:000634662800001)】;

基金:We would like to thank the reviewers for their useful comments and suggestions which helped us to considerably improve the work. We also kindly thank Ju Gao from Shuguang Hospital Affiliated to Shanghai University of Traditional Chinese Medicine for providing us clinical datasets, and Ping He from Shanghai Hospital Development Center for her help. This work was supported by the Zhejiang Lab (No. 2019ND0AB01), the National Natural Science Foundation of China (No. 61903144) and the National Key R&D Program of China for "Precision medical research" (No. 2018YFC0910550)

语种:英文

外文关键词:Active learning; Chinese word segmentation; Deep learning; Electronic health records

摘要:Electronic health records (EHRs) in hospital information systems contain patients' diagnoses and treatments, so EHRs are essential to clinical data mining. Of all the tasks in the mining process, Chinese word segmentation (CWS) is a fundamental and important one, and most state-of-the-art methods greatly rely on large scale of manually annotated data. Since annotation is time-consuming and expensive, efforts have been devoted to techniques, such as active learning, to locate the most informative samples for modeling. In this paper, we follow the trend and present an active learning method for CWS in EHRs. Specifically, a new sampling strategy combining normalized entropy with loss prediction (NE-LP) is proposed to select the most valuable data. Meanwhile, to minimize the computational cost of learning, we propose a joint model including a word segmenter and a loss prediction model. Furthermore, to capture interactions between adjacent characters, bigram features are also applied in the joint model. To illustrate the effectiveness of NE-LP, we conducted experiments on EHRs collected from the Shuguang Hospital Affiliated to Shanghai University of Traditional Chinese Medicine. The results demonstrate that NE-LP consistently outperforms conventional uncertainty-based sampling strategies for active learning in CWS

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心