Building A Large Collection of Multi-domain Electronic Theses and Dissertations

Published in IEEE International Conference on Big Data (Big Data), 2021

Recommended citation: Sami Uddin, Bipasha Banerjee, Jian Wu, William A. Ingram, and Edward A. Fox. 2021. Building A Large Collection of Multi-domain Electronic Theses and Dissertations. In 2021 IEEE International Conference on Big Data (Big Data), Or- lando, FL, USA, December 15-18, 2021. IEEE, 6043–6045. https://doi.org/10.1109/ BigData52589.2021.9672058 https://doi.org/10.1109/BigData52589.2021.9672058

In this work, we report our progress on building a collection containing over 450k Electronic Theses and Dissertations (ETDs), including full-text and metadata. Our goal is to close the gap of accessibility between long text and short text documents, and to create a new research opportunity for the scholarly community. For that, we developed an ETD Ingestion Framework (EIF) that automatically harvests metadata and PDFs of ETDs from university libraries. We faced multiple challenges and learned many lessons during the process, that led to proposed solutions to overcome/mitigate the limitations of the current data. We also described the data that we have collected. We hope our methods will be useful for building similar collections from university libraries and that the data can be used for research and education.