Explaining RL Decisions with Trajectories
Shripad Vilasrao Deshmukh, Arpan Dasgupta, Balaji Krishnamurthy, Nan Jiang, Chirag Agarwal, Georgios Theocharous, Jayakumar Subramanian
Abstract
Explanation is a key component for the adoption of reinforcement learning (RL) in many real-world decision-making problems. In the literature, the explanation is often provided by saliency attribution to the features of the RL agent's state. In this work, we propose a complementary approach to these explanations, particularly for offline RL, where we attribute the policy decisions of a trained RL agent to the trajectories encountered by it during training. To do so, we encode trajectories in offline training data individually as well as collectively (encoding a set of trajectories). We then attribute policy decisions to a set of trajectories in this encoded space by estimating the sensitivity of the decision with respect to that set. Further, we demonstrate the effectiveness of the proposed approach in terms of quality of attributions as well as practical scalability in diverse environments that involve both discrete and continuous state and action spaces such as grid-worlds, video games (Atari) and continuous control (MuJoCo). We also conduct a human study on a simple navigation task to observe how their understanding of the task compares with data attributed for a trained RL policy. Keywords -- Explainable AI, Verifiability of AI Decisions, Explainable RL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d33e2ec9-4ccc-4dd4-bdf1-c437aa2d56daCited by top-tier papers3
- A Snapshot of Influence: A Local Data Attribution Framework for Online Reinforcement LearningYuzheng Hu, Fan Wu, Haotian Ye, David A. Forsyth et al.NeurIPS 2025 · 13 citations
- SkillTree: Explainable Skill-Based Deep Reinforcement Learning for Long-Horizon Control TasksYongyan Wen, Siyuan Li, Rongchang Zuo, Lei Yuan et al.AAAI 2025 · 4 citations
- Interpreting Emergent Planning in Model-Free Reinforcement LearningThomas Bush, Stephen Chung, Usman Anwar, Adrià Garriga-Alonso et al.ICLR 2025
Builds on11
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
Related papers
- Approximating Shapley Explanations in Reinforcement LearningDaniel Beechey, Özgür SimsekNeurIPS 2025 · 1 citation
- Widening the Pipeline in Human-Guided Reinforcement Learning with Explanation and Context-Aware Data AugmentationLin Guan, Mudit Verma, Sihang Guo, Ruohan Zhang et al.NeurIPS 2021 · 57 citations
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature AttributionNikaash Puri, Sukriti Verma, Piyush Gupta, Dhruv Kayastha et al.ICLR 2020 · 99 citations
- Learning from Good Trajectories in Offline Multi-Agent Reinforcement LearningQi Tian, Kun Kuang, Furui Liu, Baoxiang WangAAAI 2023 · 14 citations
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement LearningAkanksha Atrey, Kaleigh Clary, David D. JensenICLR 2020 · 108 citations
