会议专题

ANNOTATION OF COMPLEX NOUN PHRASES FROM MULTILINGUAL PARALLEL CORPUS

  The Noun Phrase (NP) is the dominant construct in natural language text.While base NPs (BNP) and maximal length NPs (MNP) are relatively easy to identified and extracted,the internal structure of NPs is rather a challenge in natural language processing.Penn Treebank leaves the BNPs fiat as implicit right branching.Vadas and Curran added BNP internal structure to the Penn Treebank.But the results of the BNP structure are very often incorrect when it is considered within a longer complex NP (CNP).Structural ambiguity prevails in most CNPs and multilingual comparison may help improve disambiguation.We introduce a new NP annotation scheme,which is applicable to multilingual parallel corpora and discriminate genuine flat branching and right branching.Flat branching is preferred instead of binary branching wherever appropriate so as to achieve inter-lingual consistency.As a pilot task to build a gold standard corpus for structural and semantic analysis of CNPs,381 document titles are extracted from the UN resolutions as typical examples of CNPs.Document titles in Chinese,English and Russian are manually annotated in XML format with the hope to help acquire rules for parsers or machine translators targeted at CNPs.The problems encountered are reported.

Complex NPs Structural ambiguity Annotation Multilingual parallel corpus

Jingxiang Cao Degen Huang

School of Computer Science and Technology,Dalian University of Technology;School of Foreign Language School of Computer Science and Technology,Dalian University of Technology

国际会议

2012 2nd IEEE International Conference on Cloud Computing and Intelligence Systems (2012年第2届IEEE云计算与智能系统国际会议(IEEE CCIS2012))

杭州

英文

1922-1926

2012-10-30(万方平台首次上网日期,不代表论文的发表时间)