AI Interpretability RFP
Funds mechanistic interpretability research on detecting and steering deceptive behaviors in artificial intelligence systems.
The 2026 AI Interpretability RFP, issued by Schmidt Sciences under its Science of Trustworthy AI program family, funded research on detecting and correcting deceptive behaviors in AI systems. Schmidt Sciences defined deceptive behaviors to include factually incorrect statements, misleading confidence claims, fabrications, selective omission, evasiveness, and false claims regarding a model's self-k…
Mechanistic interpretability research on detecting and steering deceptive behaviors in AI models, and translating those methods into practical human-AI collaboration and multi-agent applications.
Sign up free to see the funding breakdown
Sign up free to see the industries in scope
Sign up free to see the full eligibility
Sign up free to see how to apply
Sign up free to see the timeline