详细信息
文献类型:期刊文献
中文题名:中文电子病历文本中的时间识别算法研究
英文题名:Algorithm of Time Recognition in Chinese Electronic Medical Record
作者:孙健[1];高大启[1];刘珉[2];高炬[2];阮彤[1]
机构:[1]华东理工大学信息科学与工程学院,上海200237;[2]上海曙光医院,上海200021
年份:2018
卷号:41
期号:1
起止页码:15
中文期刊名:山西大学学报(自然科学版)
外文期刊名:Journal of Shanxi University(Natural Science Edition)
收录:CSTPCD;;北大核心:【北大核心2017】;CSCD:【CSCD_E2017_2018】;
基金:国家高技术研究发展计划("863"计划)(2015AA020107);科技部科技支撑项目(2015BAH12F01-05)
语种:中文
中文关键词:独立时间;基于事件的时间;bootstrapping算法;条件随机场;命名实体识别
外文关键词::independent time; event-based time; bootstrapping algorithm; conditional random field; namedentity recognition
摘要:时间作为电子病历中的一类重要实体,对于标识患者从入院到出院期间不同阶段的病情变化,有着不可替代的作用。电子病历文本中的时间可分为独立时间和基于事件的时间,针对这两类时间分别提出了基于bootstrapping的识别算法和基于条件随机场的识别算法。其中,为了解决基于事件的时间短语太长而不能准确定位其边界的问题,引入了中文症状知识库作为词典特征,有效地提高了条件随机场识别结果的准确率、召回率和F1值。实验结果表明,该方法在独立时间和基于事件的时间识别上的F1值分别达到了92.57%和93.98%。
As an important entity in electronic medical records,the time plays an irreplaceable role on the i- dentification of changes in patients' conditions at different stages from admission to discharge. The time in electronic medical records is divided into independent time and event-based time. The two recognition algo- rithms based on bootstrapping and conditional random fields are introduced for these two types of time re- spectively. In order to solve the problem that the event-based time phrase is too long to locate its boundary accurately, a knowledge base of symptoms in Chinese is introduced as a dictionary feature, which effective- ly improves the precision, recall rate and F1 value of the recognition result of conditional random field. The experimental results show that the Flscores of the two algorithms on independent time and event-based time recognitionare 92.57 % and 93.98% respectively.
参考文献:
正在载入数据...
