Improving Word Embeddings for Antonym Detection Using Thesauri and SentiWordNet

摘要：

　　Word embedding is a distributed representation of words in a vector space.It involves a mathematical embedding from a space with one dimension per word to a continuous vector space with much lower dimension.It performs well on tasks including synonym and hyponym detection by grouping similar words.However,most existing word embeddings are insensitive to antonyms,since they are trained based on word distributions in a large amount of text data,where antonyms usually have similar contexts.To generate word embeddings that are capable of detecting antonyms,we firstly modify the objective function of Skip-Gram model,and then utilize the supervised synonym and antonym information in thesauri as well as the sentiment information of each word in SentiWordNet.We conduct evaluations on three relevant tasks,namely GRE antonym detection,word similarity,and semantic textual similarity.The experiment results show that our antonym-sensitive embedding outperforms common word embeddings in these tasks,demonstrating the efficacy of our methods.

关键词： Antonym detection Word embedding Thesauri SentiWordNet

作者: Zehao Dou Wei Wei Xiaojun Wan

作者单位: Peking University,Beijing,China

会议类型: 国际会议

会议名称: 2018自然语言处理与中文计算国际会议(NLPCC2018)

会议地点: 呼和浩特

会议语种:英文

页码: 67-79

在线出版日期: 2018-08-26（万方平台首次上网日期，不代表论文的发表时间）

会议专题

Improving Word Embeddings for Antonym Detection Using Thesauri and SentiWordNet