高级工程师
性别: 男
毕业院校: 大连理工大学
学位: 博士
所在单位: 计算机科学与技术学院
学科: 计算机应用技术
办公地点: 创新园大厦D0103房间
联系方式: QQ:2407849530
电子邮箱: xukan@dlut.edu.cn
qq : 2407849530
开通时间: ..
最后更新时间: ..
点击次数:
论文类型: 期刊论文
发表时间: 2022-06-29
发表刊物: 山东大学学报 理学版
卷号: 52
期号: 7
页面范围: 66-72
ISSN号: 1671-9352
摘要: Short text clustering plays an important role in data mining. The traditional short text clustering model has some problems, such as high dimensionality、sparse data and lack of semantic information. To overcome the shortcomings of short text clustering caused by sparse features、semantic ambiguity、dynamics and other reasons, this paper presents a feature based on the word embeddings representation of text and short text clustering algorithm based on the moving distance of the characteristic words. Initially, the word embeddings that represents semantics of the feature word was gained through training in large-scale corpus with the Continous Skip-gram Model. Furthermore, use the Euclidean distance calculation feature word similarity. Additionally, EMD (Earth Mover's Distance) was used to calculate the similarity between the short text. Finally, apply the similarity between the short text to Kmeans clustering algorithm implemented in the short text clustering. The evaluation results on three data sets show that the effect of this method is superior to traditional clustering algorithms.
备注: 新增回溯数据