详细信息

Towards anti-forgetting with masked optimal transport regularization for continual named entity recognition  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Towards anti-forgetting with masked optimal transport regularization for continual named entity recognition

作者:Ma, Zhiyuan[1];Gu, Miaomiao[1];Wang, Nan[2];Cao, Jialin[1]

机构:[1]Univ Shanghai Sci & Technol, 580 Jungong Rd, Shanghai 200093, Peoples R China;[2]East China Univ Sci & Technol, 130 Meilong Rd, Shanghai 200093, Peoples R China

年份:2026

卷号:682

外文期刊名:NEUROCOMPUTING

收录:;EI(收录号:20261320381489);WOS:【SCI-EXPANDED(收录号:WOS:001733980600003)】;

基金:Acknowledgements The work is supported by National Local Joint Engineering Laboratory of Next Generation Internet Data Processing Technology under grant No. ZYGX2025K00802. The authors would also like to thank the anonymous reviewers for their valuable comments and helpful suggestions.

语种:英文

外文关键词:Continual learining; Optimal transport; Named entity recognition; Catastrophic forgetting; Deep learning

摘要:Continual Named Entity Recognition (CNER) aims to sequentially adapt the model to emerging entities while preserving knowledge of previously learned ones. There is a common and major challenge called catastrophic forgetting in the task, which arises from two inherent issues: 1) Backward incompatibility, where entities learned in previous tasks lack annotations and are consequently labeled as "O" (non-entities) in new data; and 2) Semantic shift of the "O" label, where tokens correctly classified as non-entities under the old model may become new en tity types under the new model. Although existing methods have succeeded in mitigating catastrophic forgetting to a certain extent, they primarily address the issue through knowledge distillation, overlooking a key opportu nity, i.e. leveraging feature-space distribution discrepancies between the old and new models. To bridge this gap, this paper proposes Masked Optimal Transport Regularization (MOTR) which measures the discrepancies, thus improving the overall performance. Specifically, MOTR designs a cost function through the joint distribution of features and labels, followed by a dual-OT mechanism, which comprises an entity dealignment and align ment subtasks. The former maximizes the global probability distance between the old "O" label and the new entities to enhance discriminability for new classes, while the latter minimizes the global probability distance between consistent old/new labels to consolidate prior knowledge. By quantifying and leveraging feature-level distribution shifts via OT, MOTR effectively preserves prior knowledge while enhancing discriminability for new entities. Experiments on three publicly available benchmarks demonstrate that MOTR achieves competitive per formance, with improvements of up to 2.38% in Mi-F1 and 4.66% in Ma-F1 over other baseline methods. Notably, MOTR functions as a plug-in module compatible with multiple CNER frameworks for further performance gains, broadening methodological scope for CNER.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心