Figure Extraction from Scanned Electronic Theses and Dissertations.

Published in VtechWorks: Viginia Tech ETD, 2020

Recommended citation: Sampanna Yashwant Kahu. 2020. Figure Extraction from Scanned Electronic Theses and Dissertations. Thesis. Virginia Tech. https://vtechworks.lib.vt.edu/handle/ 10919/100113 http://hdl.handle.net/10919/100113

The ability to extract figures and tables from scientific documents can solve key use-cases such as their semantic parsing, summarization, or indexing. Although a few methods have been developed to extract figures and tables from scientific documents, their performance on scanned counterparts is considerably lower than on born-digital ones. To facilitate this, we propose methods to effectively extract figures and tables from Electronic Theses and Dissertations (ETDs), that out-perform existing methods by a considerable margin. Our contribution towards this goal is three-fold. (a) We propose a system/model for improving the performance of existing methods on scanned scientific documents for figure and table extraction. (b) We release a new dataset containing 10,182 labelled page-images spanning across 70 scanned ETDs with 3.3k manually annotated bounding boxes for figures and tables. (c) Lastly, we release our entire code and the trained model weights to enable further research (https://github.com/SampannaKahu/deepfigures-open).