详细信息

Harvesting more answer spans from paragraph beyond annotation  ( EI收录)  

文献类型:期刊文献

英文题名:Harvesting more answer spans from paragraph beyond annotation

作者:Bao, Qiaoben[1]; Chen, Jiangjie[1]; Liu, Linfang[1]; Liu, Jingping[3]; Liang, Jiaqing[1]; Xiao, Yanghua[1,2]

机构:[1] Shanghai Key Laboratory of Data Science, School of Computer Science, Fudan University, Shanghai, China; [2] Fudan-Aishu Cognitive Intelligence Joint Research Center, Shanghai, China; [3] School of Information Science and Engineering, East China University of Science and Technology, Shanghai, China

年份:2022

起止页码:27

外文期刊名:WSDM 2022 - Proceedings of the 15th ACM International Conference on Web Search and Data Mining

收录:EI(收录号:20221011761527)

基金:We thank anonymous reviewers from current and past versions of the manuscript for their comments and suggestions. This work was supported by National Key Research and Development Project (No.2020AAA0109302), Shanghai Science and Technology Innovation Action Plan (No.19511120400) and Shanghai Municipal Science and Technology Major Project (No.2021SHZDZX0103).

语种:英文

外文关键词:Natural language processing systems - Information retrieval

摘要:AutomaticA nswer spanE xtraction (AE) focuses on identifying key information from paragraphs that can be asked. It has been used to facilitate downstream question generation tasks or data augmentation for question answering. Current work of AE heavily relies on the annotated answer spans fromM achineR eadingC omprehension (MRC) datasets. However, these methods suffer from the partial annotation problem due to the annotation protocols of MRC tasks. To tackle this problem, we propose \mymethod, a S tructured Co ntext graph network with P ositive -unlabeled learning. \mymethod first represents the paragraph by constructing a graph with both syntactic and semantic edges, then adopts a unified pointer network for answer span identification. \mymethod narrows the discrenpency between AE and MRC by formulating AE as aP ositive-\textitu nlabeled (PU) learning problem, thus recovering more answer spans from paragraphs. To evaluate newly extracted spans without annotation, we also present an automatic metric from the perspective of question answering and text summarization, which correlates well with human judgments. Comprehensive experiments on both AE and downstream tasks demonstrate the effectiveness of our proposed framework. Our code is available at \urlhttps://github.com/iambabao/SCOPE. ? 2022 ACM.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心