大连理工大学主页平台管理系统 Huang Degen Incorporating Prior Knowledge into Word Embedding for Chinese Word Similarity Measurement Home

Current position: Home >> Scientific Research >> Paper Publications

Incorporating Prior Knowledge into Word Embedding for Chinese Word Similarity Measurement

Release Time:2019-03-12 Hits:

Indexed by: Journal Article

Date of Publication: 2018-05-01

Journal: ACM TRANSACTIONS ON ASIAN AND LOW-RESOURCE LANGUAGE INFORMATION PROCESSING

Included Journals: SCIE

Volume: 17

Issue: 3

ISSN: 2375-4699

Key Words: Chinese word similarity; word embedding; prior knowledge

Abstract: Word embedding-based methods have received increasing attention for their flexibility and effectiveness in many natural language-processing (NLP) tasks, including Word Similarity (WS). However, these approaches rely on high-quality corpus and neglect prior knowledge. Lexicon-based methods concentrate on human's intelligence contained in semantic resources, e.g., Tongyici Cilin, HowNet, and Chinese WordNet, but they have the drawback of being unable to deal with unknown words. This article proposes a three-stage framework for measuring the Chinese word similarity by incorporating prior knowledge obtained from lexicons and statistics into word embedding: in the first stage, we utilize retrieval techniques to crawl the contexts of word pairs from web resources to extend context corpus. In the next stage, we investigate three types of single similarity measurements, including lexicon similarities, statistical similarities, and embedding-based similarities. Finally, we exploit simple combination strategies with math operations and the counter-fitting combination strategy using optimization method. To demonstrate our system's efficiency, comparable experiments are conducted on the PKU-500 dataset. Our final results are 0.561/0.516 of Spearman/Pearson rank correlation coefficient, which outperform the state-of-the-art performance to the best of our knowledge. Experiment results on Chinese MC-30 and SemEval-2012 datasets show that our system also performs well on other Chinese datasets, which proves its transferability. Besides, our system is not language-specific and can be applied to other languages, e.g., English.

Prev One:Multi-Level Attention Based BLSTM Neural Network for Biomedical Event Extraction

Next One:基于λ-主动学习方法的中文微博分词

Home

Scientific Research

Teaching Research

Awards and Honours

Enrollment Information

Student Information

My Album

Blog

Incorporating Prior Knowledge into Word Embedding for Chinese Word Similarity Measurement