详细信息
SLR: A million-scale comprehensive crossword dataset for simultaneous learning and reasoning ( SCI-EXPANDED收录 EI收录)
文献类型:期刊文献
英文题名:SLR: A million-scale comprehensive crossword dataset for simultaneous learning and reasoning
作者:Wang, Chao[1,2];Zhu, Tinghui[3];Li, Zhixu[3];Liu, Jingping[4]
机构:[1]Shanghai Univ, Sch Future Technol, Shanghai, Peoples R China;[2]Shanghai Univ, Inst Artificial Intelligence, Shanghai, Peoples R China;[3]Fudan Univ, Sch Comp Sci, Shanghai Key Lab Data Sci, Shanghai, Peoples R China;[4]East China Univ Sci & Technol, Sch Informat Sci & Engn, Shanghai, Peoples R China
年份:2023
卷号:554
外文期刊名:NEUROCOMPUTING
收录:;EI(收录号:20233614695706);WOS:【SCI-EXPANDED(收录号:WOS:001052303400001)】;
基金:This work was supported by the Program of Natural Science Foun-dation of Shanghai (No. 23ZR1422800) .
语种:英文
外文关键词:Crossword puzzle; Knowledge reasoning; Language model; Open-domain question answering
摘要:The progress of the natural language understanding (NLU) community has put forward higher demands for the knowledge reserve and reasoning ability of the model. However, existing schemes for improving model capabilities often split them into two separate tasks. The task of enabling models to learn and reason about knowledge simultaneously has not received sufficient attention. In this paper, we propose a novel crossword-based NLU task that imparts knowledge information to a model by solving crossword clues and simultaneously trains the model to infer new knowledge from existing knowledge. To this end, we construct a comprehensive crossword dataset SLR containing more than 4 million unique clue-answer pairs. Compared to existing crossword datasets, SLR is more comprehensive and contains linguistic knowledge, expertise in various fields, and commonsense knowledge. Meanwhile, to evaluate the reasoning ability of the model, we design clever details in the reasoning of the answers, most clues require the solver to reason through two or more pieces of knowledge to arrive at an answer. We analyze the composition of the dataset and the similarities and differences of the various types of clues via sampling and consider various data partitioning methods to enhance the generalization ability of the training set. Furthermore, we test the performance of several different advanced models and methods on this dataset and analyze the strengths and weaknesses of each. An interesting conclusion is that even powerful language models perform poorly on tasks that require reasoning.
参考文献:
正在载入数据...
