详细信息

A Context-Enhanced Transformer with Abbr-Recover Policy for Chinese Abbreviation Prediction  ( EI收录)  

文献类型:期刊文献

英文题名:A Context-Enhanced Transformer with Abbr-Recover Policy for Chinese Abbreviation Prediction

作者:Cao, Kaiyan[1]; Yang, Deqing[1]; Liu, Jingping[2]; Liang, Jiaqing[1]; Xiao, Yanghua[3]; Wei, Feng[4]; Wu, Baohua[4]; Lu, Quan[4]

机构:[1] School of Data Science, Fudan University, Shanghai, China; [2] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China; [3] School of Computer Science, Fudan University, Shanghai, China; [4] Alibaba Group, Hangzhou, China

年份:2022

起止页码:2944

外文期刊名:International Conference on Information and Knowledge Management, Proceedings

收录:EI(收录号:20224413038297)

语种:英文

外文关键词:Forecasting - Iterative decoding - Natural language processing systems - Semantics

摘要:Chinese abbreviation prediction is very important for various natural language processing tasks such as query understanding and entity linking, since people tend to use the concise abbreviation rather than the full form (name) to mention an entity. The existing models achieve their predictions through sequence labeling, i.e., the binary classification for each character (token) of the full form. However, they only leverage the semantics of the entity itself, overlooking the label dependencies between the tokens, and the rich information of the entity-related texts. In this paper we proposed a Context-Enhanced Transformer with Abbr-Recover policy, namely CETAR, for Chinese abbreviation prediction. CETAR predicts the abbreviation sequence mainly through an iterative decoding process, of which each round consists of an abbreviation and recovery operation. Our extensive experiments upon both general field and specific domain datasets justify that CETAR outperforms the state-of-the-art baselines including sequence labeling models and sequence generation models. Moreover, we have successfully constructed a Chinese abbreviation dataset from the famous tour website Fliggy, and we also shared it at https://github.com/tolerancecky/abbr-0731. The online A/B test on the Fliggy search system shows that 2.03% of conversion rate improvement has been achieved with the predicted abbreviations. ? 2022 ACM.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心