TripleTree: A Versatile Interpretable Representation of Black Box Agents and their Environments
Tom Bewley, Jonathan Lawry
Abstract
In explainable artificial intelligence, there is increasing interest in understanding the behaviour of autonomous agents to build trust and validate performance. Modern agent architectures, such as those trained by deep reinforcement learning, are currently so lacking in interpretable structure as to effectively be black boxes, but insights may still be gained from an external, behaviourist perspective. Inspired by conceptual spaces theory, we suggest that a versatile first step towards general understanding is to discretise the state space into convex regions, jointly capturing similarities over the agent's action, value function and temporal dynamics within a dataset of observations. We create such a representation using a novel variant of the CART decision tree algorithm, and demonstrate how it facilitates practical understanding of black box agents through prediction, visualisation and rule-based explanation. * Supported by an EPSRC/Thales industrial CASE award. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f6cdf5d-8cba-4b15-8cf5-0002e6ee35a2Cited by top-tier papers12
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 79 citations
- StateMask: Explaining Deep Reinforcement Learning through State MaskZelei Cheng, Xian Wu, Jiahao Yu, Wenhai Sun et al.NeurIPS 2023 · 24 citations
- Learning Tree Interpretation from Object Representation for Deep Reinforcement LearningGuiliang Liu, Xiangyu Sun, Oliver Schulte, Pascal PoupartNeurIPS 2021 · 16 citations
- Accountability in Offline Reinforcement Learning: Explaining Decisions with a Corpus of ExamplesHao Sun, Alihan Hüyük, Daniel Jarrett, Mihaela van der SchaarNeurIPS 2023 · 13 citations
- Understanding Individual Agent Importance in Multi-Agent System via Counterfactual ReasoningJianming Chen, Yawen Wang, Junjie Wang, Xiaofei Xie et al.AAAI 2025 · 11 citations
Related papers
- Local Explanations for Reinforcement LearningRonny Luss, Amit Dhurandhar, Miao LiuAAAI 2023 · 5 citations
- ProtoX: Explaining a Reinforcement Learning Agent via PrototypingRonilo J. Ragodos, Tong Wang, Qihang Lin, Xun ZhouNeurIPS 2022 · 13 citations
- Asking the Right Questions: Learning Interpretable Action Models Through Query AnsweringPulkit Verma, Shashank Rao Marpally, Siddharth SrivastavaAAAI 2021 · 33 citations
- This State Looks Like That: Self-Interpretable Reinforcement Learning Agents using Prototype Soft Actor-CriticAndrea Marzo, Alessio Ragno, Roberto CapobiancoICML 2026
- SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control TasksYongyan Wen, Siyuan Li, Rongchang Zuo, Lei Yuan et al.AAAI 2025 · 4 citations
