详细信息

基于关键词相似度的短文本分类方法研究    

Research on short text classification based on keyword similarity

文献类型:期刊文献

中文题名:基于关键词相似度的短文本分类方法研究

英文题名:Research on short text classification based on keyword similarity

作者:张振豪[1];过弋[1,2,3];韩美琪[1];王吉祥[1]

机构:[1]华东理工大学信息科学与工程学院,上海200237;[2]石河子大学信息科学与技术学院,新疆石河子832003;[3]大数据流通与交易技术国家工程实验室--商业智能与可视化技术研究中心,上海200436

年份:2020

卷号:37

期号:1

起止页码:26

中文期刊名:计算机应用研究

外文期刊名:Application Research of Computers

收录:CSTPCD;;北大核心:【北大核心2017】;CSCD:【CSCD_E2019_2020】;

基金:国家自然科学基金资助项目(61462073);上海市科学技术委员会项目(17DZ1101003,18511106602).

语种:中文

中文关键词:词向量;特征选择;短文本分类;特征权重

外文关键词:word embedding;feature selecting;short text classification;feature weighting

摘要:在传统的文本分类中,文本向量空间矩阵存在维数灾难和极度稀疏等问题,而提取与类别最相关的关键词作为文本分类的特征有助于解决以上两个问题。针对以上结论进行研究,提出了一种基于关键词相似度的短文本分类框架。该框架首先通过大量语料训练得到word2vec词向量模型;然后通过TextRank获得每一类文本的关键词,在关键词集合中进行去重操作作为特征集合。对于任意特征,通过词向量模型计算短文本中每个词与该特征的相似度,选择最大相似度作为该特征的权重。最后选择K近邻(KNN)和支持向量机(SVM)作为分类器训练算法。实验基于中文新闻标题数据集,与传统的短文本分类方法相比,分类效果约平均提升了6%,从而验证了该框架的有效性。
In order to cope with the problem of data sparsity and curse of dimensionality in text classification,this paper proposed a short text classification framework by taking keyword as features and assigning keyword similarity as feature weight.First,it trained a word2 vec model with large corpus data,then got keywords of each category text by textrank.And it selected unique keywords from the keywords collection as features.For each feature,it calculated the similarity of words in the short text by word2 vec model,and assigned the maximum similarity as the weight of the feature.Finally,it chose KNN and SVM as classifier.Experiments on dataset of Chinese news headlines demonstrate that the accuracy outperforms other usual methods by 6%.

参考文献:

正在载入数据...

版权所有©华东理工大学 重庆维普资讯有限公司 渝B2-20050021-7 
渝公网安备 50019002500408号 违法和不良信息举报中心