Shields to Guarantee Probabilistic Safety in MDPs
Linus Heck, Filip Macák, Roman Andriushchenko, Milan Ceska, Sebastian Junges
Abstract
Abstract Shielding is a prominent model-based technique to ensure safety of autonomous agents. Classical shielding aims to ensure that nothing bad ever happens and comes with strong guarantees about safety and maximal permissiveness. However, shielding systems for probabilistic safety, where something bad is allowed to happen with an acceptable probability, has proven to be more intricate. This paper presents a formal framework that conservatively extends classical shields to probabilistic safety. In this framework, we (i) demonstrate the impossibility of preserving the strong guarantees on safety and permissiveness, (ii) provide natural shields with weaker guarantees, and (iii) introduce offline and online shield constructions ensuring strong safety guarantees. The empirical evaluation highlights the practical advantages of the new shields, as well as their computational feasibility.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21ab5a7d-d096-4a2c-8964-4996e670cfc9Builds on4
- Shield Synthesis for LTL Modulo TheoriesAndoni Rodríguez, Guy Amir, Davide Corsi, César Sánchez et al.AAAI 2025 · 13 citations
- Probabilistic Shielding for Safe Reinforcement LearningEdwin Hamel-De le Court, Francesco Belardinelli, Alexander W. GoodallAAAI 2025 · 7 citations
- Reinforcement Learning of Risk-Constrained Policies in Markov Decision ProcessesTomás Brázdil, Krishnendu Chatterjee, Petr Novotný, Jiri VahalaAAAI 2020 · 5 citations
- Adaptive Shielding via Parametric Safety ProofsYao Feng, Jun Zhu, André Platzer, Jonathan LaurentOOPSLA 2025 · 4 citations
Related papers
- Explainably Safe Reinforcement LearningSabine Rieder, Stefan Pranger, Debraj Chakraborty, Jan Kretínský et al.NeurIPS 2025
- Safe Exploration in Reinforcement Learning by Reachability Analysis over Learned ModelsYuning Wang, He ZhuCAV 2024 · 2 citations
- Guiding Safe Exploration with Weakest PreconditionsGreg Anderson, Swarat Chaudhuri, Isil DilligICLR 2023
- Classification with Conceptual SafeguardsHailey Joren, Charles T. Marx, Berk UstunICLR 2024 · 3 citations
- Robust Adaptive Multi-Step Predictive ShieldingTanmay Ambadkar, Darshan Chudiwal, Greg Anderson, Abhinav VermaICLR 2026
