This project is a collaboration with Harvard University and the Astrophysics Data System (ADS), a digital library portal operated by the Smithsonian Astrophysical Observatory (SAO) under a NASA grant. With over 15 million records, ADS is one of the most important archives in the scientific field of astronomy.

"Newer documents are ‘born digital,’ making them machine-readable and parseable," said Naiman. "This has not only helped domain scientists find relevant research more efficiently, but through methods like natural language processing, it also has facilitated new discoveries in these fields."

Naiman's project aims to extend these capabilities to predigital documents by extracting their text, figures, and tables, allowing researchers to apply the same information mining methods that are available to "born digital" documents. This will result in more easily searchable documents and new discoveries. The work will also enhance the screen-reading capabilities of these documents to make them more accessible.

