How RL Agents Behave When Their Actions Are Modified
Eric D. Langlois, Tom Everitt
Abstract
Reinforcement learning in complex environments may require supervision to prevent the agent from attempting dangerous actions. As a result of supervisor intervention, the executed action may differ from the action specified by the policy. How does this affect learning? We present the Modified-Action Markov Decision Process, an extension of the MDP model that allows actions to differ from the policy. We analyze the asymptotic behaviours of common reinforcement learning algorithms in this setting and show that they adapt in different ways: some completely ignore modifications while others go to various lengths in trying to avoid action modifications that decrease reward. By choosing the right algorithm, developers can prevent their agents from learning to circumvent interruptions or constraints, and better control agent responses to other kinds of action modification, like self-damage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 52fb3f96-be44-4e7c-837a-34370eacbf22Cited by top-tier papers4
- Agent Incentives: A Causal PerspectiveTom Everitt, Ryan Carey, Eric D. Langlois, Pedro A. Ortega et al.AAAI 2021 · 66 citations
- Towards Safe Reinforcement Learning with a Safety Editor PolicyHaonan Yu, Wei Xu, Haichao ZhangNeurIPS 2022 · 50 citations
- A Complete Criterion for Value of Information in Soluble Influence DiagramsChris van Merwijk, Ryan Carey, Tom EverittAAAI 2022 · 7 citations
- Learning "Partner-Aware" Collaborators in Multi-Party CollaborationAbhijnan Nath, Nikhil KrishnaswamyNeurIPS 2025 · 2 citations
Related papers
- High Confidence Generalization for Reinforcement LearningJames E. Kostas, Yash Chandak, Scott M. Jordan, Georgios Theocharous et al.ICML 2021 · 5 citations
- Acting in Delayed Environments with Non-Stationary Markov PoliciesEsther Derman, Gal Dalal, Shie MannorICLR 2021 · 4 citations
- Safe Reinforcement Learning via Curriculum InductionMatteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause et al.NeurIPS 2020 · 109 citations
- Learn to change the world: Multi-level reinforcement learning with model-changing actionsZiqing Lu, Babak Hassibi, Lifeng Lai, Weiyu XuICML 2026
- Average-Reward Learning and Planning with OptionsYi Wan, Abhishek Naik, Richard S. SuttonNeurIPS 2021 · 12 citations
