会议专题

Improving Word Embeddings for Antonym Detection Using Thesauri and SentiWordNet

  Word embedding is a distributed representation of words in a vector space.It involves a mathematical embedding from a space with one dimension per word to a continuous vector space with much lower dimension.It performs well on tasks including synonym and hyponym detection by grouping similar words.However,most existing word embeddings are insensitive to antonyms,since they are trained based on word distributions in a large amount of text data,where antonyms usually have similar contexts.To generate word embeddings that are capable of detecting antonyms,we firstly modify the objective function of Skip-Gram model,and then utilize the supervised synonym and antonym information in thesauri as well as the sentiment information of each word in SentiWordNet.We conduct evaluations on three relevant tasks,namely GRE antonym detection,word similarity,and semantic textual similarity.The experiment results show that our antonym-sensitive embedding outperforms common word embeddings in these tasks,demonstrating the efficacy of our methods.

Antonym detection Word embedding Thesauri SentiWordNet

Zehao Dou Wei Wei Xiaojun Wan

Peking University,Beijing,China

国际会议

2018自然语言处理与中文计算国际会议(NLPCC2018)

呼和浩特

英文

67-79

2018-08-26(万方平台首次上网日期,不代表论文的发表时间)