Saute RL: Almost Surely Safe Reinforcement Learning Using State Augmentation
Aivar Sootla, Alexander I. Cowen-Rivers, Taher Jafferjee, Ziyan Wang, David Henry Mguni, Jun Wang, Haitham Ammar
Abstract
Satisfying safety constraints almost surely (or with probability one) can be critical for the deployment of Reinforcement Learning (RL) in reallife applications. For example, plane landing and take-off should ideally occur with probability one. We address the problem by introducing Safety Augmented (Sauté) Markov Decision Processes (MDPs), where the safety constraints are eliminated by augmenting them into the state-space and reshaping the objective. We show that Sauté MDP satisfies the Bellman equation and moves us closer to solving Safe RL with constraints satisfied almost surely. We argue that Sauté MDP allows viewing the Safe RL problem from a different perspective enabling new features. For instance, our approach has a plug-and-play nature, i.e., any RL algorithm can be "Sautéed". Additionally, state augmentation allows for policy generalization across safety constraints. We finally show that Sauté RL algorithms can outperform their state-of-the-art counterparts when constraint satisfaction is of high importance. 1 Huawei R&D UK.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03df7264-add3-42b6-ad2b-8da24e87e13aCited by top-tier papers5
- Compositional Policy Learning in Stochastic Control Systems with Formal GuaranteesDorde Zikelic, Mathias Lechner, Abhinav Verma, Krishnendu Chatterjee et al.NeurIPS 2023 · 31 citations
- Probabilistic Shielding for Safe Reinforcement LearningEdwin Hamel-De le Court, Francesco Belardinelli, Alexander W. GoodallAAAI 2025 · 7 citations
- MANSA: Learning Fast and Slow in Multi-Agent SystemsDavid Henry Mguni, Haojun Chen, Taher Jafferjee, Jianhong Wang et al.ICML 2023 · 4 citations
- Finding Safe Zones of Markov Decision Processes PoliciesLee Cohen, Yishay Mansour, Michal MoshkovitzNeurIPS 2023 · 1 citation
- Learning Robust Multi-Agent Policies via Selective Adversarial Fault InductionDavid H Mguni, Yaqi Sun, Haojun Chen, Wanrong Yang et al.ICML 2026
Builds on6
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable ModelAlex X. Lee, Anusha Nagabandi, Pieter Abbeel, Sergey LevineNeurIPS 2020 · 437 citations
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 252 citations
- Safe Reinforcement Learning via Curriculum InductionMatteo Turchetta, Andrey Kolobov, Shital Shah, Andreas Krause et al.NeurIPS 2020 · 109 citations
Related papers
- Safety Representations for Safer Policy LearningKaustubh Mani, Vincent Mai, Charlie Gauthier, Annie S. Chen et al.ICLR 2025
- Safe Reinforcement Learning Using Advantage-Based InterventionNolan Wagener, Byron Boots, Ching-An ChengICML 2021 · 66 citations
- Constrained Markov Decision Processes via Backward Value FunctionsHarsh Satija, Philip Amortila, Joelle PineauICML 2020 · 58 citations
- DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement LearningArchana Bura, Aria HasanzadeZonuzy, Dileep Kalathil, Srinivas Shakkottai et al.NeurIPS 2022 · 48 citations
- Near-Optimal Sample Complexity for Online Constrained MDPsChang Liu, Yunfan Li, Lin F. YangNeurIPS 2025 · 1 citation
