详细信息
Keywords extraction based on sentence-ranking from Chinese patents ( EI收录)
文献类型:期刊文献
英文题名:Keywords extraction based on sentence-ranking from Chinese patents
作者:Wang, Zhihong[1]; Guo, Yi[1,2]; Qi, Tianmei[1]
机构:[1] Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, China; [2] School of Information Science and Technology, Shihezi University, Xinjiang, China
年份:2018
卷号:2018-July
起止页码:80
外文期刊名:Proceedings of the International Conference on Software Engineering and Knowledge Engineering, SEKE
收录:EI(收录号:20184706125793)
语种:英文
外文关键词:Patents and inventions - Knowledge engineering - Software engineering - Information retrieval systems - Extraction
摘要:Patent, an important scientific literature, records a large amount of innovative and practical research. The patent keywords also provide a high-level topic description of a patent document and hold an important position in classic NLP tasks, such as patent classification or clustering. However, there are few research works on keywords extraction covering the Chinese patents in current stage. In this paper, we propose a novel algorithm to extract keywords from Chinese patents. A sentenceranking model, based on a sentence embedding graph and heuristic rules, is constructed to select the top-KS percent of the sentences. At the same time, the semantic-ranking weights of sentences are also transmitted to keywords extraction. The experimental results on our Chinese patents datasets testifies that the sentence-ranking based keywords extraction algorithm improves the performance by 6% to 13% in F-score. In summary, the new idea of selecting key sentences from original documents can effectively filter out noisy sentences and leverage the performance of keywords extraction. ? 2018 Universitat zu Koln. All rights reserved.
参考文献:
正在载入数据...
