![]() |
个人信息Personal Information
副教授
硕士生导师
主要任职:Director of Practice Center of Foreign Languages
其他任职:语言与认知研究所副所长
性别:女
毕业院校:大连理工大学
学位:博士
所在单位:外国语学院
学科:外国语言学及应用语言学. 计算机应用技术
办公地点:海晏楼701
联系方式:caojx@dlut.edu.cn
电子邮箱:caojx@dlut.edu.cn
ANNOTATION OF COMPLEX NOUN PHRASES FROM MULTILINGUAL PARALLEL CORPUS
点击次数:
论文类型:会议论文
发表时间:2012-01-01
收录刊物:CPCI-S
页面范围:1440-1444
关键字:Complex NPs; Structural ambiguity; Annotation; Multilingual parallel corpus
摘要:The Noun Phrase (NP) is the dominant construct in natural language text. While base NPs (BNP) and maximal length NPs (MNP) are relatively easy to identified and extracted, the internal structure of NPs is rather a challenge in natural language processing. Penn Treebank leaves the BNPs flat as implicit right branching. Vadas and Curran added BNP internal structure to the Penn Treebank. But the results of the BNP structure are very often incorrect when it is considered within a longer complex NP (CNP). Structural ambiguity prevails in most CNPs and multilingual comparison may help improve disambiguation. We introduce a new NP annotation scheme, which is applicable to multilingual parallel corpora and discriminate genuine flat branching and right branching. Flat branching is preferred instead of binary branching wherever appropriate so as to achieve inter-lingual consistency. As a pilot task to build a gold standard corpus for structural and semantic analysis of CNPs, 381 document titles are extracted from the UN resolutions as typical examples of CNPs. Document titles in Chinese, English and Russian are manually annotated in XML format with the hope to help acquire rules for parsers or machine translators targeted at CNPs. The problems encountered are reported.