Iterative Bounding MDPs: Learning Interpretable Policies via Non-Interpretable Methods
Nicholay Topin, Stephanie Milani, Fei Fang, Manuela Veloso
摘要
Current work in explainable reinforcement learning generally produces policies in the form of a decision tree over the state space. Such policies can be used for formal safety verification, agent behavior prediction, and manual inspection of important features. However, existing approaches fit a decision tree after training or use a custom learning procedure which is not compatible with new learning techniques, such as those which use neural networks. To address this limitation, we propose a novel Markov Decision Process (MDP) type for learning decision tree policies: Iterative Bounding MDPs (IBMDPs). An IBMDP is constructed around a base MDP so each IBMDP policy is guaranteed to correspond to a decision tree policy for the base MDP when using a method-agnostic masking procedure. Because of this decision tree equivalence, any function approximator can be used during training, including a neural network, while yielding a decision tree policy for the base MDP. We present the required masking procedure as well as a modified value update step which allows IBMDPs to be solved using existing algorithms. We apply this procedure to produce IBMDP variants of recent reinforcement learning methods. We empirically show the benefits of our approach by solving IBMDPs to produce decision tree policies for the base MDPs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- End-to-End Neuro-Symbolic Reinforcement Learning with Textual ExplanationsLirui Luo, Guoxi Zhang, Hongming Xu, Yaodong Yang 等ICML 2024 · 被引用 18 次
- A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement LearningYuzheng Hu, Fan Wu, Haotian Ye, David A. Forsyth 等NeurIPS 2025 · 被引用 13 次
- Small Decision Trees for MDPs with Deductive SynthesisRoman Andriushchenko, Milan Ceska, Sebastian Junges, Filip MacákCAV 2025 · 被引用 2 次
- Breiman meets Bellman: Non-Greedy Decision Trees with MDPsHector Kohler, Riad Akrour, Philippe PreuxKDD 2025
- LICORICE: Label-Efficient Concept-Based Interpretable Reinforcement LearningZhuorui Ye, Stephanie Milani, Geoffrey J. Gordon, Fei FangICLR 2025
相关 Paper
- SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control TasksYongyan Wen, Siyuan Li, Rongchang Zuo, Lei Yuan 等AAAI 2025 · 被引用 4 次
- Explainably Safe Reinforcement LearningSabine Rieder, Stefan Pranger, Debraj Chakraborty, Jan Kretínský 等NeurIPS 2025
- Safety Verification of Decision-Tree Policies in Continuous TimeChristian Schilling, Anna Lukina, Emir Demirovic, Kim Guldstrand LarsenNeurIPS 2023 · 被引用 5 次
- SPOT: Scalable Policy Optimization with Trees for Markov Decision ProcessesXuyuan Xiong, Pedro Chumpitaz-Flores, Kaixun Hua, Cheng HuaNeurIPS 2025
- RGMDT: Return-Gap-Minimizing Decision Tree Extraction in Non-Euclidean Metric SpaceJingdi Chen, Hanhan Zhou, Yongsheng Mei, Carlee Joe-Wong 等NeurIPS 2024 · 被引用 2 次
