Explainably Safe Reinforcement Learning
Sabine Rieder, Stefan Pranger, Debraj Chakraborty, Jan Kretínský, Bettina Könighofer
Abstract
Trust in a decision-making system requires both safety guarantees and the ability to interpret and understand its behavior. This is particularly important for learned systems, whose decision-making processes are often highly opaque. Shielding is a prominent model-based technique for enforcing safety in reinforcement learning. However, because shields are automatically synthesized using rigorous formal methods, their decisions are often similarly difficult for humans to interpret. Recently, decision trees became customary to represent controllers and policies. However, since shields are inherently non-deterministic, their decision tree representations become too large to be explainable in practice. To address this challenge, we propose a novel approach for explainable safe RL that enhances trust by providing human-interpretable explanations of the shield's decisions. Our method represents the shielding policy as a hierarchy of decision trees, offering top-down, case-based explanations. At design time, we use a world model to analyze the safety risks of executing actions in given states. Based on this risk analysis, we construct both the shield and a high-level decision tree that classifies states into risk categories (safe, critical, dangerous, unsafe), providing an initial explanation of why a given situation may be safety-critical. At runtime, we generate localized decision trees that explain which actions are allowed and why others are deemed unsafe. Altogether, our method facilitates the explainability of the safety aspect in the safe-by-shielding reinforcement learning. Our framework requires no additional information beyond what is already used for shielding, incurs minimal overhead, and can be readily integrated into existing shielded RL pipelines. In our experiments, we compute explanations using decision trees that are several orders of magnitude smaller than the original shield.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18bde87c-5e5b-4e0d-8ab4-6dbcb652ba90Builds on6
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 79 citations
- Shield Decentralization for Safe Multi-Agent Reinforcement LearningDaniel Melcer, Christopher Amato, Stavros TripakisNeurIPS 2022 · 26 citations
- StateMask: Explaining Deep Reinforcement Learning through State MaskZelei Cheng, Xian Wu, Jiahao Yu, Wenhai Sun et al.NeurIPS 2023 · 24 citations
- Test Where Decisions Matter: Importance-driven Testing for Deep Reinforcement LearningStefan Pranger, Hana Chockler, Martin Tappler, Bettina KönighoferNeurIPS 2024 · 7 citations
- Small Decision Trees for MDPs with Deductive SynthesisRoman Andriushchenko, Milan Ceska, Sebastian Junges, Filip MacákCAV 2025 · 2 citations
Related papers
- Shield Synthesis for LTL Modulo TheoriesAndoni Rodríguez, Guy Amir, Davide Corsi, César Sánchez et al.AAAI 2025 · 13 citations
- Adaptive Shielding via Parametric Safety ProofsYao Feng, Jun Zhu, André Platzer, Jonathan LaurentOOPSLA 2025 · 4 citations
- Iterative Bounding MDPs: Learning Interpretable Policies via Non-Interpretable MethodsNicholay Topin, Stephanie Milani, Fei Fang, Manuela VelosoAAAI 2021 · 45 citations
- Probabilistic Shielding for Safe Reinforcement LearningEdwin Hamel-De le Court, Francesco Belardinelli, Alexander W. GoodallAAAI 2025 · 7 citations
- Shields to Guarantee Probabilistic Safety in MDPsLinus Heck, Filip Macák, Roman Andriushchenko, Milan Ceska et al.CAV 2026
