A Rule-based Framework of Metadata Extraction from Scientific Papers

摘要：

Most scientific documents on the web are unstructured or semi-structured, and the automatic document metadata extraction process becomes an important task. This paper describes a framework for automatic metadata extraction from scientific papers. Based on a spatial and visual knowledge principle, our system can extract title, authors and abstract from scientific papers. We utilize format information such as font size and position to guide the metadata extraction process. The experiment results show that our system achieves a high accuracy in header metadata extraction which can effectively assist the automatic index creation for digital libraries.

关键词： document metadata information extraction rulebased

作者: Zhixin Guo Hai Jin

作者单位: Cluster and Grid Computing Lab Services Computing Technology and System Lab Huazhong University of Science and Technology, Wuhan, 430074, China

会议类型: 国际会议

会议名称: 2011 IEEE 10th International Symposium on Distributed Computing and Applications to Business,Engineering(第十届电子商务、工程及科学领域的分布式计算和应用国际学术研讨会 DCABES 2011)

会议地点: 无锡

会议语种:英文

页码: 400-404

在线出版日期: 2011-10-14（万方平台首次上网日期，不代表论文的发表时间）

会议专题

A Rule-based Framework of Metadata Extraction from Scientific Papers