会议专题

Statistical Machine Translation Based on LDA

Current Statistical Machine Translation (SMT) systems translate one sentence at a time, ignoring any document level information. Consequently, translation models are learned only at sentence level and document contexts are generally overlooked. In this paper, we try to introduce document topic to help SMT system to produce target sentences. First, the parallel training corpus with underlying document boundary is segmented into multiple documents, and then we use a monolingual LDA model to determine which topics these documents belong to. Next, the background phrase table is enhanced with the probability distribution of a document over topics. Evaluation shows that our proposed approach significantly improves the BLEU score on Chinese-to-English machine translation.

SMT LDA Adaptation Document

Gong Zhengxian Zhang Yu Zhou Guodong

School of Computer Science and Technology Soochow University Suzhou, China

国际会议

2010 4th International Universal Communication Symposium(第四届国际普遍交流学术研讨会 IUCS 2010)

北京

英文

285-289

2010-10-18(万方平台首次上网日期,不代表论文的发表时间)