详细信息

Keywords extraction based on sentence-ranking from Chinese patents  ( EI收录)  

文献类型:期刊文献

英文题名:Keywords extraction based on sentence-ranking from Chinese patents

作者:Wang, Zhihong[1]; Guo, Yi[1,2]; Qi, Tianmei[1]

机构:[1] Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, China; [2] School of Information Science and Technology, Shihezi University, Xinjiang, China

年份:2018

卷号:2018-July

起止页码:80

外文期刊名:Proceedings of the International Conference on Software Engineering and Knowledge Engineering, SEKE

收录:EI(收录号:20184706125793)

语种:英文

外文关键词:Patents and inventions - Knowledge engineering - Software engineering - Information retrieval systems - Extraction

摘要:Patent, an important scientific literature, records a large amount of innovative and practical research. The patent keywords also provide a high-level topic description of a patent document and hold an important position in classic NLP tasks, such as patent classification or clustering. However, there are few research works on keywords extraction covering the Chinese patents in current stage. In this paper, we propose a novel algorithm to extract keywords from Chinese patents. A sentenceranking model, based on a sentence embedding graph and heuristic rules, is constructed to select the top-KS percent of the sentences. At the same time, the semantic-ranking weights of sentences are also transmitted to keywords extraction. The experimental results on our Chinese patents datasets testifies that the sentence-ranking based keywords extraction algorithm improves the performance by 6% to 13% in F-score. In summary, the new idea of selecting key sentences from original documents can effectively filter out noisy sentences and leverage the performance of keywords extraction. ? 2018 Universitat zu Koln. All rights reserved.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心