Towards Zero-Tax Alignment of Responsible Frontier LLMs via Adaptive Inference-Time Steering
Total Funding to Date
Investigator
- Dong Wang
The project aims to develop a new approach that could help AI better distinguish between requests that require intervention and those that do not. AI systems are designed to avoid harmful or biased responses. But sometimes those protections can be too broad, causing an AI to reject legitimate questions or lose important context needed to provide a good answer. The researchers would assess an AI prompt as it is being processed and determine whether an intervention is actually needed. When one is necessary, the researchers will study how to make it as targeted as possible, reducing harmful behavior without unnecessarily interfering with the AI’s ability to reason, retrieve information or answer other parts of a question. The broader goal is to help make AI systems safer while still ensuring that they are helpful.
Personnel
Funding Agencies
- Amazon, FY27 – $82,830.00