详细信息

Sentence-Ranking-Enhanced Keywords Extraction from Chinese Patents  ( SCI-EXPANDED收录 EI收录)  

文献类型:期刊文献

英文题名:Sentence-Ranking-Enhanced Keywords Extraction from Chinese Patents

作者:Wang, Zhi-Hong[1];Guo, Yi[1,2,3]

机构:[1]East China Univ Sci & Technol, Dept Comp Sci & Engn, Shanghai 200237, Peoples R China;[2]Natl Engn Lab Big Data Distribut & Exchange Techn, Business Intelligence & Visualizat Res Ctr, Shanghai 200436, Peoples R China;[3]Shihezi Univ, Sch Informat Sci & Technol, Shihezi 8320003, Peoples R China

年份:2019

卷号:35

期号:3

起止页码:651

外文期刊名:JOURNAL OF INFORMATION SCIENCE AND ENGINEERING

收录:;EI(收录号:20192006937410);WOS:【SSCI(收录号:WOS:000467782400011),SCI-EXPANDED(收录号:WOS:000467782400011)】;

基金:This research is financially supported by National Key Research and Development Program of China Grant No. 2018YFC0807105, National Natural Science Foundation of China Grant No. 61462073 and Science and Technology Committee of Shanghai Municipality (STCSM) Grant Nos. 17DZ1101003, 18511106602 and I 8DZ2252300.

语种:英文

外文关键词:Chinese patents; key sentences; sentence-ranking model; keywords extraction; human-annotated dataset

摘要:Patent keywords, a high-level topic representation of patents, hold an important position in many patent-oriented mining tasks, such as classification, retrieval and translation. However, there are few studies concentrated on keywords extraction for patents in current stage, and neither exist human-annotated gold standard datasets, especially for Chinese patents. This paper introduces a new human-annotated Chinese patent dataset and proposes a sentence-ranking based Term Frequency-Inverse Document Frequency (SR based TF-IDF) algorithm for patent keywords extraction, motivated by the thought of "the keywords are in the key sentences". In the algorithm, a sentence-ranking model is constructed to filter top-K-s percent sentences from each patent based on a sentence semantic graph and heuristic rules. At last, the proposed algorithm is evaluated with TF-IDF, TextRank, word2vec weighted TextRank and Patent Keyword Extraction Algorithm (PKEA) on the homemade Chinese patent dataset and several standard benchmark datasets. The experimental results testify that our proposed algorithm effectively improves the performance of extracting keywords from Chinese patents.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心