详细信息
文献类型:期刊文献
中文题名:基于LDA-BERT相似性测度模型的文本主题演化研究
英文题名:Topic Evolution Research Based on LDA-BERT Similarity Measure Model
作者:海骏林峰[1];严素梅[1];陈荣[1];李建霞[1]
机构:[1]华东理工大学科技信息研究所,上海200237
年份:2024
期号:1
起止页码:72
中文期刊名:图书馆工作与研究
外文期刊名:Library Work and Study
收录:;国家哲学社会科学学术期刊数据库;北大核心:【北大核心2023】;CSSCI:【CSSCI_E2023_2024】;
语种:中文
中文关键词:相似性测度;LDA-BERT模型;LDA模型;BERT模型;主题演化
外文关键词:Similarity measure;LDA-BERT model;LDA model;BERT model;Theme evolution
摘要:文章针对LDA主题模型在提取文本主题时忽略文本语义关联的问题,提出基于LDA-BERT的相似性测度模型:首先,结合利用TF-IDF和TextRank方法提取文本特征词,利用LDA主题模型挖掘文本主题;其次,通过嵌入BERT模型,结合LDA主题模型构建的主题-主题词概率分布,从词粒度层面表示主题向量;最后,利用余弦相似度算法计算主题之间的相似度。在相似性测度模型基础上构建向量相似度指标分析文献研究主题之间的关联,并绘制主题演化知识图谱。通过智慧图书馆领域的实证研究发现,使用LDA-BERT模型计算出的主题相似度结果相较于LDA主题模型的计算结果更加准确,与实际情况更相符。
Aiming at the problem that the traditional LDA topic model ignores the semantic correlation when extracting text topics,this paper proposes a similarity measure model based on LDA-BERT.Firstly,by combining TF-IDF and TextRank methods,text feature words are extracted and text topics are mined using LDA model.Secondly,by embedding BERT model and combining LDA topic model,the probability distribution of subject-subject words is constructed to represent the topic vector from the level of word granularity.Finally,cosine similarity algorithm is used to calculate the similarity between subjects.Based on the similarity measure model,the vector similarity index was constructed to analyze the correlation between literature research topics,and the knowledge map of topic evolution was drawn.The empirical research was carried out in the field of smart library.It is found that the results calculated by the LDA-BERT model are more accurate than those of the topic LDA model calculations,and more consistent with the actual situation.
参考文献:
正在载入数据...
