Assistant Professor Jiaqi Ma has received a National Science Foundation (NSF) CAREER award to develop new tools to understand how individual components of training data affect the behavior of large artificial intelligence systems. The highly selective CAREER award recognizes early-career faculty with the potential to serve as academic role models in research and education and to lead advances in the mission of their department or organization. Ma received a five-year, $660,307 grant for his project, "Data Attribution and Curation for Web-Scale AI Systems," which seeks to improve data attribution, a family of methods that estimate how training examples influence model behavior.
AI systems increasingly influence how people search for information, receive recommendations, learn, work, and create. The quality of these systems depends heavily on the data used to train them. However, modern training data is often enormous, noisy, and constantly changing, because it is drawn from many sources. Poorly understood data can reduce accuracy, amplify harmful information, weaken reasoning and/or make it difficult to recognize the value of content in AI systems.
"Data is one of the most important ingredients in modern AI," Ma said, "but we still have a limited understanding of how individual training examples shape a model's behavior."
Ma hopes to develop tools to help researchers and practitioners decide what data to keep, remove, prioritize, or compensate for. By making data curation more principled and transparent, the project has the potential to improve the performance, reliability, and safety of widely used technologies such as language models and recommendation systems.
"This project aims to make those influences measurable, even in large and continuously evolving AI systems," he said. "Better data attribution can help us build models that are more accurate, reliable, and safe, while also creating a more transparent basis for recognizing and compensating the people whose content contributes to these systems."
The project will also create public educational resources, open-source software, conference tutorials, course modules, and research opportunities for graduate, undergraduate, and pre-college students, helping widen access to data-centered AI research and training.
Ma's research interests lie in the broad area of machine learning and AI, with recent focuses on the data foundations of AI, including three complementary aspects: 1) understanding how training data impact AI models (data attribution); 2) developing data-centric algorithms that improve the quality and safety of training data (data curation and synthetic data generation); and 3) studying how data mediate the societal impact of AI (data compensation and machine unlearning). His work has been recognized with a Best Paper Award from the Workshop on Navigating and Addressing Data Problems for Foundation Models at the 2024 International Conference on Learning Representations, and a New Faculty Highlight at the Association for the Advancement of Artificial Intelligence 2025.
He earned his PhD from the University of Michigan and worked as a postdoctoral researcher at Harvard University.