Explaining Reinforcement Learning with Shapley Values
Daniel Beechey, Thomas M. S. Smith, Özgür Simsek
Abstract
For reinforcement learning systems to be widely adopted, their users must understand and trust them. We present a theoretical analysis of explaining reinforcement learning using Shapley values, following a principled approach from game theory for identifying the contribution of individual players to the outcome of a cooperative game. We call this general framework Shapley Values for Explaining Reinforcement Learning (SVERL). Our analysis exposes the limitations of earlier uses of Shapley values in reinforcement learning. We then develop an approach that uses Shapley values to explain agent performance. In a variety of domains, SVERL produces meaningful explanations that match and supplement human intuition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d152d95-5fb6-4271-a689-4a337446cbe3Cited by top-tier papers5
- Refining Diffusion Planner for Reliable Behavior Synthesis by Automatic Detection of Infeasible PlansKyowoon Lee, Seongun Kim, Jaesik ChoiNeurIPS 2023 · 31 citations
- Towards Multi-dimensional Explanation Alignment for Medical ClassificationLijie Hu, Songning Lai, Wenshuo Chen, Hongru Xiao et al.NeurIPS 2024 · 8 citations
- A Comprehensive Study of Shapley Value in Data AnalyticsHong Lin, Shixin Wan, Zhongle Xie, Ke Chen et al.VLDB 2025 · 4 citations
- Approximating Shapley Explanations in Reinforcement LearningDaniel Beechey, Özgür SimsekNeurIPS 2025 · 1 citation
- Feature Importance Metrics in the Presence of Missing DataHenrik von Kleist, Joshua Wendland, Ilya Shpitser, Carsten MarrICML 2025
Builds on2
Related papers
- Approximating the Shapley Value without Marginal ContributionsPatrick Kolpaczki, Viktor Bengs, Maximilian Muschalik, Eyke HüllermeierAAAI 2024 · 43 citations
- Causal Shapley Values: Exploiting Causal Knowledge to Explain Individual Predictions of Complex ModelsTom Heskes, Evi Sijben, Ioan Gabriel Bucur, Tom ClaassenNeurIPS 2020 · 235 citations
- Neural Payoff Machines: Predicting Fair and Stable Payoff Allocations Among Team MembersDaphne Cornelisse, Thomas Rood, Yoram Bachrach, Mateusz Malinowski et al.NeurIPS 2022 · 10 citations
- SHAQ: Incorporating Shapley Value Theory into Multi-Agent Q-LearningJianhong Wang, Yuan Zhang, Yunjie Gu, Tae-Kyun KimNeurIPS 2022 · 50 citations
- RankSHAP: Shapley Value Based Feature Attributions for Learning to RankTanya Chowdhury, Yair Zick, James AllanICLR 2025
