详细信息
A sequence-to-sequence model for large-scale chinese abbreviation database construction ( EI收录)
文献类型:期刊文献
英文题名:A sequence-to-sequence model for large-scale chinese abbreviation database construction
作者:Wang, Chao[1]; Liu, Jingping[3]; Zhuang, Tianyi[1]; Li, Jiahang[1]; Liu, Juntao[1]; Xiao, Yanghua[1,2]; Wang, Wei[1]; Xie, Rui[4]
机构:[1] Fudan University, 1Shanghai Key Laboratory of Data Science, School of Computer Science, Shanghai, China; [2] Fudan-Aishu Cognitive Intelligence Joint Research Center, Shanghai, China; [3] East China University of Science and Technology, School of Information Science and Engineering, Shanghai, China; [4] Meituan, Shanghai, China
年份:2022
起止页码:1063
外文期刊名:WSDM 2022 - Proceedings of the 15th ACM International Conference on Web Search and Data Mining
收录:EI(收录号:20221011761457)
基金:We thank anonymous reviewers from current and past versions of the manuscript for their comments and suggestions. This work was supported by National Key Research and Development Project (No.2020AAA0109302), Shanghai Science and Technology Innovation Action Plan (No.19511120400) and Shanghai Municipal Science and Technology Major Project (No.2021SHZDZX0103).
语种:英文
外文关键词:Forecasting - Natural language processing systems
摘要:Abbreviations often used in our daily communication play an important role in natural language processing. Most of the existing studies regard the Chinese abbreviation prediction as a sequence labeling problem. However, sequence labeling models usually ignore label dependencies in the process of abbreviation prediction, and the label prediction of each character should be conditioned on its previous labels. In this paper, we propose to formalize the Chinese abbreviation prediction task as a sequence generation problem, and a novel sequence-to-sequence model is designed. To boost the performance of our deep model, we further propose a multi-level pre-trained model that incorporates character, word, and concept-level embeddings. To evaluate our methods, a new dataset for Chinese abbreviation prediction is automatically built, which contains 81,351 pairs of full forms and abbreviations. Finally, we conduct extensive experiments on a public dataset and the built dataset, and the experimental results on both datasets show that our model outperforms the state-of-the-art methods. More importantly, we build a large-scale database for a specific domain, i.e., life services in Meituan Inc., with high accuracy of about 82.7%, which contains 4,134,142 pairs of full forms and abbreviations. The online A/B testing on Meituan APP and Dianping APP suggests that Click-Through Rate increases by 0.59% and 0.86% respectively when the built database is used in the searching system. We have released our API on http://kw.fudan.edu.cn/ddemos/abbr/ with over 87k API calls in 9 months. ? 2022 ACM.
参考文献:
正在载入数据...
