Learning Tree Interpretation from Object Representation for Deep Reinforcement Learning
Guiliang Liu, Xiangyu Sun, Oliver Schulte, Pascal Poupart
摘要
Interpreting Deep Reinforcement Learning (DRL) models is important to enhance trust and comply with transparency regulations. Existing methods typically explain a DRL model by visualizing the importance of low-level input features with superpixels, attentions, or saliency maps. Our approach provides an interpretation based on high-level latent object features derived from a disentangled representation. We propose a Represent And Mimic (RAMi) framework for training 1) an identifiable latent representation to capture the independent factors of variation for the objects and 2) a mimic tree that extracts the causal impact of the latent features on DRL action values. To jointly optimize both the fidelity and the simplicity of a mimic tree, we derive a novel Minimum Description Length (MDL) objective based on the Information Bottleneck (IB) principle. Based on this objective, we describe a Monte Carlo Regression Tree Search (MCRTS) algorithm that explores different splits to find the IB-optimal mimic tree. Experiments show that our mimic tree achieves strong approximation performance with significantly fewer nodes than baseline models. We demonstrate the interpretability of our mimic tree by showing latent traversals, decision rules, causal impacts, and human evaluation results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Interpreting Unfairness in Graph Neural Networks via Training Node AttributionYushun Dong, Song Wang, Jing Ma, Ninghao Liu 等AAAI 2023 · 被引用 32 次
- Local Explanations for Reinforcement LearningRonny Luss, Amit Dhurandhar, Miao LiuAAAI 2023 · 被引用 5 次
- SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control TasksYongyan Wen, Siyuan Li, Rongchang Zuo, Lei Yuan 等AAAI 2025 · 被引用 4 次
- AIRS: Explanation for Deep Reinforcement Learning based Security ApplicationsJiahao Yu, Wenbo Guo, Qi Qin, Gang Wang 等USENIX Security 2023
它引用的顶会 Paper4
- Contrastive Learning of Structured World ModelsThomas N. Kipf, Elise van der Pol, Max WellingICLR 2020 · 被引用 322 次
- TripleTree: A Versatile Interpretable Representation of Black Box Agents and their EnvironmentsTom Bewley, Jonathan LawryAAAI 2021 · 被引用 33 次
- Contrastive Explanations for Reinforcement Learning via Embedded Self PredictionsZhengxian Lin, Kin-Ho Lam, Alan FernICLR 2021 · 被引用 28 次
- Explaining by Imitating: Understanding Decisions by Interpretable Policy LearningAlihan Hüyük, Daniel Jarrett, Cem Tekin, Mihaela van der SchaarICLR 2021 · 被引用 22 次
相关 Paper
- Delphi: A Neuro-Symbolic Framework for Individualized, Safe and Interpretable Treatment RecommendationMuchan Tao, Haonan Qin, Yuqi Fang, Caifeng Shan 等AAAI 2026
- Explaining A Black-box By Using A Deep Variational Information Bottleneck ApproachSeo-Jin Bang, Pengtao Xie, Heewook Lee, Wei Wu 等AAAI 2021 · 被引用 33 次
- RGMDT: Return-Gap-Minimizing Decision Tree Extraction in Non-Euclidean Metric SpaceJingdi Chen, Hanhan Zhou, Yongsheng Mei, Carlee Joe-Wong 等NeurIPS 2024 · 被引用 2 次
- Iterative Bounding MDPs: Learning Interpretable Policies via Non-Interpretable MethodsNicholay Topin, Stephanie Milani, Fei Fang, Manuela VelosoAAAI 2021 · 被引用 45 次
- Counterfactual Concept Bottleneck ModelsGabriele Dominici, Pietro Barbiero, Francesco Giannini, Martin Gjoreski 等ICLR 2025
