School of Information Sciences

Towards Zero-Tax Alignment of Responsible Frontier LLMs via Adaptive Inference-Time Steering

Total Funding to Date

$82,830.00

Investigator

  • Dong Wang

The project aims to develop a new approach that could help AI better distinguish between requests that require intervention and those that do not. AI systems are designed to avoid harmful or biased responses. But sometimes those protections can be too broad, causing an AI to reject legitimate questions or lose important context needed to provide a good answer. The researchers would assess an AI prompt as it is being processed and determine whether an intervention is actually needed. When one is necessary, the researchers will study how to make it as targeted as possible, reducing harmful behavior without unnecessarily interfering with the AI’s ability to reason, retrieve information or answer other parts of a question. The broader goal is to help make AI systems safer while still ensuring that they are helpful. 

Photo of hands typing on keyboard with illustrations of checks and x marks to illustrate AI safety

Funding Agencies

  • Amazon, FY27 – $82,830.00

School of Information Sciences

501 E. Daniel St.

MC-493

Champaign, IL

61820-6211

Voice: (217) 333-3280

Email: ischool@illinois.edu

Back to top