Underwood receives NEH grant to investigate consequences of error in digital libraries

Ted Underwood
Ted Underwood, Professor

Professor Ted Underwood has received a $73,122 grant from the National Endowment for the Humanities to investigate the consequences of error in digital libraries. While digital libraries represent an immense storehouse of knowledge, the texts are full of errors because of the imperfect process by which they are transcribed optically.

"It isn't unusual for five percent of the words in volumes to be mistranscribed, with the level of error much higher in some volumes," said Underwood. "Simply measuring the fraction of mistranscribed words is easy. It’s harder to know how much difference those errors make for the methods and questions that actually interest researchers. Some forms of analysis are undisturbed by high levels of error; others may be quite sensitive, especially when errors are distributed unevenly across different historical periods and genres."

Underwood will work with graduate students from the iSchool and English Department to construct parallel collections that pair each "clean" text with a realistically error-ridden version of the same book drawn from a digital library. The team will build collections of Chinese texts as well as English texts ranging from 1700 to the present, because different character sets and printing technologies produce different kinds of error. Then the team will apply a wide range of data-mining methods to both the clean and error-ridden collections and measure the distortion produced by transcription error and other common sources of noise. The project will provide tools that help other researchers estimate the level of uncertainty in their own conclusions.

"No data is perfect. There's always some kind of error. The question is whether the error is of a kind and magnitude likely to matter for a particular question," he said.

Underwood is a professor in the iSchool and also holds an appointment with the Department of English in the College of Liberal Arts and Sciences. He has authored three books about literary history, including Distant Horizons (The University of Chicago Press Books, 2019), Why Literary Periods Mattered: Historical Contrast and the Prestige of English Studies (Stanford University Press, 2013), and The Work of the Sun: Literature, Science and Political Economy 1760-1860 (New York: Palgrave, 2005). His articles have appeared in PMLA, Representations, MLQ, and Cultural Analytics. Underwood earned his PhD in English from Cornell University.

Updated on
Backto the news archive

Related News

Hoiem receives Schiller Prize for “Education of Things”

Associate Professor Elizabeth Hoiem has won the 2025 Justin G. Schiller Prize from The Bibliographical Society of America for her book, The Education of Things: Mechanical Literacy in British Children's Literature, 1762-1860 (University of Massachusetts Press). The prize, which recognizes the best bibliographical work on pre-1951 children's literature, includes a cash award of $3,000 and a year's membership in the Society. 

Elizabeth Hoiem

Chan authors new book connecting eugenics and Big Tech

Associate Professor Anita Say Chan has authored a new book that identifies how the eugenics movement foreshadows the predatory data tactics used in today's tech industry. Her book, Predatory Data: Eugenics in Big Tech and Our Fight for an Independent Future, was released this month by the University of California Press and featured in the news outlets San Francisco Chronicle and Mother Jones.

Anita Say Chan

CCB contributes to new Books to Parks site on Lyddie

The Center for Children's Books (CCB) collaborated with the National Park Service (NPS) to launch a new Books to Parks website on Lyddie, a 1991 novel by Katherine Paterson that highlights the experiences of young women working in textile mills in nineteenth-century Lowell, Massachusetts. 

Lyddie book

Layne-Worthey edits book on digital humanities and LIS

Glen Layne-Worthey, associate director for research support services for the HathiTrust Research Center (HTRC), and Isabel Galina, researcher at the Institute for Bibliographic Studies at the National University of Mexico, have edited a new book, The Routledge Companion to Libraries, Archives, and the Digital Humanities, which was recently released by Routledge.

Glen Layne-Worthey