Corrigibility Transformation: Constructing Goals That Accept Updates
Rubi Hudson
Abstract
An AI agent will learn a desired goal more effectively if it does not resist the training process, but many partially learned goals incentivize an AI to avoid further goal updates. We would like goals to be corrigible, meaning they allow changes requested through designated channels, so that we can confidently correct errors and shut down the AI if necessary. Despite this being a crucial safety property, the existing literature does not specify goals that are both corrigible and competitive with alternatives. We introduce a transformation that constructs a corrigible version of nearly any goal, without sacrificing performance. This is done by eliciting predictions of reward conditional on costlessly preventing updates, and having that target be pursued myopically. These goals are then shown to lead to optimal performance among the class of corrigible goals, to incentivize allowing mid-action overrides, and to disincentivize deliberate self-modification. Empirically, they induce corrigible behavior in gridworld settings and for language models when applied at the prompt level.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb00833c-2aba-4cf2-8a8c-c6cb9e27a4afBuilds on5
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIsMantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim et al.NeurIPS 2025 · 84 citations
- ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test CasesZiqian Zhong, Aditi Raghunathan, Nicholas CarliniICLR 2026 · 54 citations
- AI Alignment with Changing and Influenceable Reward FunctionsMicah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell et al.ICML 2024 · 44 citations
- Reinforcement Learning in Newcomblike EnvironmentsJames Bell, Linda Linsefors, Caspar Oesterheld, Joar SkalseNeurIPS 2021 · 23 citations
- Why Do Some Language Models Fake Alignment While Others Don't?Abhay Sheshadri, John Hughes, Julian Michael, Alex Mallen et al.NeurIPS 2025 · 19 citations
Related papers
- Recontextualization Mitigates Specification Gaming Without Modifying the SpecificationAriana Azarbal, Victor Gillioz, Vladimir Ivanov, Bryce Woodworth et al.ICML 2026 · 11 citations
- Rule Based Rewards for Language Model SafetyTong Mu, Alec Helyar, Johannes Heidecke, Joshua Achiam et al.NeurIPS 2024 · 159 citations
- MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward HackingSebastian Farquhar, Vikrant Varma, David Lindner, David K. Elson et al.ICML 2025
- Output Supervision Can Obfuscate the Chain of ThoughtJacob Drori, Luke Marks, Bryce Woodworth, Alex Cloud et al.ICLR 2026 · 10 citations
- Learning "Partner-Aware" Collaborators in Multi-Party CollaborationAbhijnan Nath, Nikhil KrishnaswamyNeurIPS 2025 · 2 citations
