详细信息

基于BERT预训练模型的事故案例文本分类方法    

Text Classification Method of Accident Cases Based on BERT Pre-Training Model

文献类型:期刊文献

中文题名:基于BERT预训练模型的事故案例文本分类方法

英文题名:Text Classification Method of Accident Cases Based on BERT Pre-Training Model

作者:涂远来[1];周家乐[1];王慧锋[1]

机构:[1]华东理工大学信息科学与工程学院,上海200237

年份:2023

卷号:49

期号:4

起止页码:576

中文期刊名:华东理工大学学报(自然科学版)

外文期刊名:Journal of East China University of Science and Technology

收录:Scopus;北大核心:【北大核心2020】;CSCD:【CSCD_E2023_2024】;

基金:青年科学基金(61906068);国家重点研发计划(2018YFC1803306)。

语种:中文

中文关键词:危险辨识;文本分类;BERT;需求分析;安全攸关系统

外文关键词:hazard identification;text classification;BERT;requirement analysis;safety critical system

摘要:事故案例数据库中的大量事故信息为安全攸关系统的设计提供了丰富、宝贵的经验,包括事故发生的时间、地点、原因、经过等。这些信息在危险辨识中起着至关重要的作用,但它们通常分布在事故文档的各个段落中,使得人工提取的效率低且成本高。本文提出了一种基于BERT(Bidirectional Encoder Representations from Transformers)预训练模型的事故案例文本分类方法,可将事故案例文本分为ACCIDENT、CAUSE、CONSEQUENCE、RESPONSE这4类。此外,收集并构建了事故案例文本数据集用于训练模型。实验结果表明,本文方法可以实现对事故案例文本的自动分类,分类准确率达到73.44%,召回率为69.13%,F1值为0.71。
The large amount of accident information in the accident case database can provide rich and valuable experience for the design of safety related system,including time,location,cause,process of accidents,etc.These informations play an important role in hazard identification,but they are usually distributed in various paragraphs of accident documents,which makes manual extraction inefficient and costly.This paper proposes a text classification method for accident cases based on BERT pre-training model,which can classify accident case texts into four categories:ACCIDENT、CAUSE、CONSEQUENCE,and RESPONSE.In addition,a test dataset of accident cases is collected and produced for training the model.The experiment shows that this method can achieve the automatic classification of accident case text,with a classification accuracy of 73.44%,a recall rate of 69.13%,and an F1 value of 0.71.In this paper,multiple groups of different experimental parameters are set up,and the effect of parameter settings on classification is fully explored through experiments to find the best parameter settings.The proposed classification method can help better mine the semantic information in the accident case text and provide powerful technical support for the subsequent establishment of expert knowledge base and efficient accident retrieval platform.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心