School of Information Sciences

Is code enough? Stodden, Marinov research focuses on providing code rather than generated data with research

Sharing and reusing research data is becoming increasingly common in the scientific world, allowing researchers to more easily build on the work of others as they seek new discoveries.

A new research project being conducted by Associate Professor Victoria Stodden and Illinois Computer Science Professor Darko Marinov aims to answer key questions about how researchers can reliably share the code used to generate their data rather than the more costly data itself.

"Our question was, when is it possible to save only the code that produced simulated data, and that’s all I need to save, and when do I also need to save the data?" Stodden said. "Simulation codes can produce massive amounts of data, for example petabytes of data. If I can rerun the code and regenerate the data, in theory I don't even need to save the data. For what types of codes is that possible? That's exactly the question we're trying to answer."

The National Science Foundation is funding the work by Stodden and Marinov, providing $300,000 over two years. Stodden, who also is an Illinois Computer Science faculty affiliate, is the PI on the project. Marinov is the co-PI.

As Stodden, who is the lead investigator for the project, explains, the format for scholarly articles has changed little in decades. It provides only a small space to discuss how researchers derived their results.

But as computation has become more integral to research across virtually every scientific field, that format has become inadequate, she said.

"There's such an amount of complexity – the computer can do X calculations per second. So how do you actually explain the increased complexity of computational research in words in a small section in a paper? It can be very, very difficult," Stodden said.

Now some journals, she said, have begun to require researchers to publish their data and code along with their findings.

Stodden and Marinov, an expert on the testing and reliability of software, wondered whether providing the code alone could reliably allow the results of a given paper to be reproduced. And if it is the code that accompanies the published research, what kind of standard should it meet?

"If code is going to travel with this scholarly output, the community will need to come to some type agreement regarding code standards," she said.

For their project, the two are focusing on physics research as an example because of its intensive computational needs.

In preliminary work using articles from the Journal of Computational Physics, Stodden and her group tried to replicate the computational results from 55 articles and were unable to reproduce any. After contacting the authors, Stodden says they came away with the impression that many believed reproducing their computational results would be straightforward, something they found not to be the case.

Eventually, Stodden and Marinov hope to determine whether and how code could be reliably substituted for data for a wide range of fields.

"We want to learn how to do better scientific software, software that is more reliable, and that researchers can trust more," Stodden said. "These questions have come about not because the scientific community isn't doing a good job; they came about because computation is so important, and increasingly so. We're chasing fascinating opportunities here."

Research Areas:
Tags:
Updated on
Backto the news archive

Related News

Dahlen selected as juror for 2026 Kirkus Prize

Associate Professor Sarah Park Dahlen has been selected as one of six jurors for the 2026 Kirkus Prize, given annually in the categories of fiction, nonfiction, and young readers' literature. The prize is one of the richest in the literary world, with awards of $50,000 in each category.

Sarah Park Dahlen

Liu receives support for AI project through NVIDIA Academic Grant Program

Assistant Professor Yaoyao Liu has been awarded a grant through the NVIDIA Academic Grant Program. NVIDIA, a world leader in accelerated computing and AI, established the program to advance academic research by providing world-class computing access and resources to researchers. Liu has received 32,000 A100 GPU-hours on Brev, an AI and machine learning platform that empowers developers to run, build, train, deploy, and scale AI models with GPU in the cloud. 

Yaoyao Liu

New app designed to improve conference experience

A new app developed by Associate Professor Yun Huang aims to make navigating conferences less work and more fun, so that attendees can meet others, discover fresh ideas, and "experience academic life as an exciting adventure." The app, PapersClaw.fun, will debut at the ACM Conference on Human Factors in Computing Systems (CHI 2026), which will be held from April 13-17 in Barcelona, Spain.

Yun Huang

Seo selected as CAS Beckman Fellow

Assistant Professor JooYoung Seo has been selected as a Center for Advanced Study (CAS) Beckman Fellow for the 2026-2027 academic year. CAS is one of the most prestigious faculty recognition programs at the University of Illinois. Its primary mission is to identify and support the most productive and innovative faculty across all disciplines. CAS Fellows are nominated by their unit heads and selected by the Center's permanent faculty through a competitive review process, with final approval by the Board of Trustees. 

JooYoung Seo

iSchool participation in iConference 2026

The following iSchool faculty and students will participate in iConference 2026, which will be held virtually from March 23–26 and physically from March 29–April 2 in Edinburgh, Scotland. The theme of this year's conference is "Information Literacies, Authenticity and Use: The Move Towards a Digitally Enlightened Society."

School of Information Sciences

501 E. Daniel St.

MC-493

Champaign, IL

61820-6211

Voice: (217) 333-3280

Email: ischool@illinois.edu

Back to top