详细信息

Text Classification Model Enhanced by Unlabeled Data for LaTeX Formula  ( SCI-EXPANDED收录)  

文献类型:期刊文献

英文题名:Text Classification Model Enhanced by Unlabeled Data for LaTeX Formula

作者:Cheng, Hua[1];Yu, Renjie[1];Tang, Yixin[1];Fang, Yiquan[1];Cheng, Tao[1]

机构:[1]East China Univ Sci & Technol, Sch Informat Sci & Engn, 130 Meilong Rd, Shanghai 200237, Peoples R China

年份:2021

卷号:11

期号:22

外文期刊名:APPLIED SCIENCES-BASEL

收录:;WOS:【SCI-EXPANDED(收录号:WOS:000725445000001)】;

语种:英文

外文关键词:unlabeled data; self-training; pretraining; BERT; LaTeX formula

摘要:Generic language models pretrained on large unspecific domains are currently the foundation of NLP. Labeled data are limited in most model training due to the cost of manual annotation, especially in domains including massive Proper Nouns such as mathematics and biology, where it affects the accuracy and robustness of model prediction. However, directly applying a generic language model on a specific domain does not work well. This paper introduces a BERT-based text classification model enhanced by unlabeled data (UL-BERT) in the LaTeX formula domain. A two-stage Pretraining model based on BERT(TP-BERT) is pretrained by unlabeled data in the LaTeX formula domain. A double-prediction pseudo-labeling (DPP) method is introduced to obtain high confidence pseudo-labels for unlabeled data by self-training. Moreover, a multi-rounds teacher-student model training approach is proposed for UL-BERT model training with few labeled data and more unlabeled data with pseudo-labels. Experiments on the classification of the LaTex formula domain show that the classification accuracies have been significantly improved by UL-BERT where the F1 score has been mostly enhanced by 2.76%, and lower resources are needed in model training. It is concluded that our method may be applicable to other specific domains with enormous unlabeled data and limited labelled data.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心