Explaining Reinforcement Learning Agents through Counterfactual Action Outcomes
Yotam Amitai, Yael Septon, Ofra Amir
Abstract
Explainable reinforcement learning (XRL) methods aim to help elucidate agent policies and decision-making processes. The majority of XRL approaches focus on local explanations, seeking to shed light on the reasons an agent acts the way it does at a specific world state. While such explanations are both useful and necessary, they typically do not portray the outcomes of the agent's selected choice of action. In this work, we propose ``COViz'', a new local explanation method that visually compares the outcome of an agent's chosen action to a counterfactual one. In contrast to most local explanations that provide state-limited observations of the agent's motivation, our method depicts alternative trajectories the agent could have taken from the given state and their outcomes. We evaluated the usefulness of COViz in supporting people's understanding of agents' preferences and compare it with reward decomposition, a local explanation method that describes an agent's expected utility for different actions by decomposing it into meaningful reward types. Furthermore, we examine the complementary benefits of integrating both methods. Our results show that such integration significantly improved participants' performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54901196-e9b2-427b-be78-3fa61214ec70Builds on4
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature AttributionNikaash Puri, Sukriti Verma, Piyush Gupta, Dhruv Kayastha et al.ICLR 2020 · 99 citations
- Contrastive Explanations for Reinforcement Learning via Embedded Self PredictionsZhengxian Lin, Kin-Ho Lam, Alan FernICLR 2021 · 28 citations
- "I Don't Think So": Summarizing Policy Disagreements for Agent ComparisonYotam Amitai, Ofra AmirAAAI 2022 · 13 citations
Related papers
- What Did You Think Would Happen? Explaining Agent Behaviour through Intended OutcomesHerman Yau, Chris Russell, Simon HadfieldNeurIPS 2020 · 44 citations
- ProtoX: Explaining a Reinforcement Learning Agent via PrototypingRonilo J. Ragodos, Tong Wang, Qihang Lin, Xun ZhouNeurIPS 2022 · 13 citations
- Counterfactual Effect Decomposition in Multi-Agent Sequential Decision MakingStelios Triantafyllou, Aleksa Sukovic, Yasaman Zolfimoselo, Goran RadanovicICML 2025
- UTILITY: Utilizing Explainable Reinforcement Learning to Improve Reinforcement LearningShicheng Liu, Minghui ZhuICLR 2025
- CrystalBox: Future-Based Explanations for Input-Driven Deep RL SystemsSagar Patel, Sangeetha Abdu Jyothi, Nina NarodytskaAAAI 2024 · 1 citation
