School of Information Sciences

HathiTrust Research Center receives NEH support for open research tools

Stephen Downie
J. Stephen Downie, Professor, Executive Associate Dean, and Co-Director of the HathiTrust Research Center

The HathiTrust Research Center (HTRC), cohosted by the iSchool at Illinois and the Luddy School of Informatics at Indiana University, has received a $325,000 Digital Humanities Advancement Grant from the National Endowment for the Humanities. One of 15 awarded nationwide, this grant will support the development of a new set of visualizations, analytical tools, and infrastructure to enable users to interact more directly with the rich data extracted from the HathiTrust Digital Library’s collection of more than 17.5 million digitized volumes.

The project, "Tools for Open Research and Computation with HathiTrust: Leveraging Intelligent Text Extraction" (TORCHLITE) will be led jointly by Professor J. Stephen Downie, associate dean for research at the iSchool and co-director of HTRC, and John Walsh, HTRC director and associate professor of information and library science at Indiana University. HTRC staff at both universities will collaborate during the next two years to accomplish TORCHLITE's goals.

"TORCHLITE will enable us to increase dramatically, and to open more fully, public access to the massive, rich data that HTRC has created from the HathiTrust Digital Library corpus," said Downie. "We have already developed innovative ways to transform, enhance, and provide access to data created from the many millions of scanned books held by HathiTrust. With TORCHLITE, we'll create new methods for accessing this data, together with several easy-to-use tools to allow people to interact with it, analyze it, and visualize it in novel ways."

The data of interest is contained in HTRC's flagship "Extracted Features" (EF) dataset, which consists of rich metadata and statistical information inferred by algorithm from the digitized texts of the entire HathiTrust corpus and documents every word on every page, including the number of times the word appears, its part of speech, and other formal features of the language on the page. The EF dataset, and methods for computing over it, have enabled many forms of full-text analysis—even of copyrighted materials. The EF dataset contains nearly three trillion tokens (or in other words, words) representing more than six billion pages of text, making it arguably the largest open dataset of its kind that is readily available to researchers around the world.

In addition to creating new methods for dealing with this enormous data set, and perhaps more impactfully, Downie emphasized that TORCHLITE will develop a framework on which digital humanities scholars, digital librarians, data scientists, and anyone else interested in textual analysis, can build and implement their own tools, and deploy them in their own environments, in order to access the HTRC data more directly and openly. TORCHLITE's tools and methods will enable the retrieval of standard volume-level descriptors—such as title, publisher, date of publication, genre, and page count—along with page- and word-level linguistic and statistical information.

In addition to creating interactive, easy-to-use tools and dashboards, TORCHLITE will promote broad community engagement through a workshop and a mentored hackathon in autumn 2023, in the hopes of encouraging individual researchers to develop their own tools using the project’s application programming interface (API).

HTRC is the official research arm of HathiTrust, a library consortium that hosts books owned and digitized by its member libraries, often in cooperation with the Google Books Project and other mass-digitization efforts. Its mission is to contribute to the common good by collecting, organizing, preserving, communicating, and sharing the record of human knowledge through data-intensive computational methods.

Updated on
Backto the news archive

Related News

iSchool launches Summer Intensive

This summer, iSchool students will have the opportunity to enroll in select courses through the new Summer Intensive pilot program, which will take place on campus over the course of two weeks. Each course will run for one week, with lessons lasting all day. Students may enroll in courses for one or both weeks, for a maximum of four credit hours. In addition to the all-day classes, students will enjoy a range of academic, professional, and social events in the evenings and on the adjoining weekends.

Aerial view of Illinois

He inducted into Sigma Xi

Professor Jingrui He has been inducted into Sigma Xi, The Scientific Research Honor Society. Sigma Xi is the international honor society of science and engineering and one of the oldest and largest scientific organizations in the world, boasting a history of service to science and society spanning over 125 years. It has a multidisciplinary membership of scientists, engineers, and scholars, and Sigma Xi chapters can be found in universities and colleges, government laboratories, and commercial research centers.

Jingrui He

Hassan and Bashir receive distinguished paper award

A paper co-authored by PhD student Muhammad Hassan and Associate Professor Masooda Bashir received the Distinguished Paper Award at the Workshop on Security and Privacy in Standardized IoT, which was held last month in San Diego, California, in conjunction with the Network and Distributed System Security (NDSS) Symposium 2026. 

iSchool researchers to present work at Technocracy Conference

This week, iSchool PhD students and faculty will present their research at the Technocracy Conference. Hosted by the Unit for Criticism and Interpretive Theory at the University of Illinois on March 5–6, the conference will begin with a panel of graduate student papers and continue the following day with invited speakers and a keynote. All events will take place at the Levis Faculty Center on the Urbana campus. 

Fab Lab summer camps foster creativity and hands-on learning

With topics like printmaking, weaving, and Minecraft 3D, it isn't surprising that summer camps offered by the Champaign-Urbana (CU) Community Fab Lab fill up so quickly. Throughout seven weeks this summer, the Fab Lab, a makerspace that supports campus and public community members, will hold 26 week-long camps for youth aged 10 to 15. This summer marks the tenth anniversary of the Fab Lab summer camps.

A camper participates in printmaking during summer camp at the Champaign-Urbana Community Fab Lab.

School of Information Sciences

501 E. Daniel St.

MC-493

Champaign, IL

61820-6211

Voice: (217) 333-3280

Email: ischool@illinois.edu

Back to top