Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable Representations
Sarath Sreedharan, Utkarsh Soni, Mudit Verma, Siddharth Srivastava, Subbarao Kambhampati
Abstract
As increasingly complex AI systems are introduced into our daily lives, it becomes important for such systems to be capable of explaining the rationale for their decisions and allowing users to contest these decisions. A significant hurdle to allowing for such explanatory dialogue could be the vocabulary mismatch between the user and the AI system. This paper introduces methods for providing contrastive explanations in terms of user-specified concepts for sequential decision-making settings where the system's model of the task may be best represented as an inscrutable model. We do this by building partial symbolic models of a local approximation of the task that can be leveraged to answer the user queries. We test these methods on a popular Atari game (Montezuma's Revenge) and variants of Sokoban (a well-known planning benchmark) and report the results of user studies to evaluate whether people find explanations generated in this form useful.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4a1b8cf-64b6-4f73-835d-349dc86a16baCited by top-tier papers7
- State2Explanation: Concept-Based Explanations to Benefit Agent Learning and User UnderstandingDevleena Das, Sonia Chernova, Been KimNeurIPS 2023 · 33 citations
- Goal Alignment: Re-analyzing Value Alignment Problems Using Human-Aware AIMalek Mechergui, Sarath SreedharanAAAI 2024 · 18 citations
- Autonomous Capability Assessment of Sequential Decision-Making Systems in Stochastic SettingsPulkit Verma, Rushang Karia, Siddharth SrivastavaNeurIPS 2023 · 14 citations
- Optimistic Exploration in Reinforcement Learning Using Symbolic Model EstimatesSarath Sreedharan, Michael KatzNeurIPS 2023 · 12 citations
- Explaining Decentralized Multi-Agent Reinforcement Learning PoliciesKayla Boggess, Sarit Kraus, Lu FengAAAI 2026
Builds on6
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
- Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement LearningAkanksha Atrey, Kaleigh Clary, David D. JensenICLR 2020 · 108 citations
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature AttributionNikaash Puri, Sukriti Verma, Piyush Gupta, Dhruv Kayastha et al.ICLR 2020 · 99 citations
- Machine versus Human Attention in Deep Reinforcement Learning TasksSihang Guo, Ruohan Zhang, Bo Liu, Yifeng Zhu et al.NeurIPS 2021 · 38 citations
Related papers
- Contrastive Explanations That Anticipate Human Misconceptions Can Improve Human Decision-Making SkillsZana Buçinca, Siddharth Swaroop, Amanda E. Paluch, Finale Doshi-Velez et al.CHI 2025 · 31 citations
- Learning Global Transparent Models consistent with Local Contrastive ExplanationsTejaswini Pedapati, Avinash Balakrishnan, Karthikeyan Shanmugam, Amit DhurandharNeurIPS 2020 · 35 citations
- Learning to Explain Selectively: A Case Study on Question AnsweringShi Feng, Jordan L. Boyd-GraberEMNLP 2022 · 4 citations
- What I Cannot Predict, I Do Not Understand: A Human-Centered Evaluation Framework for Explainability MethodsJulien Colin, Thomas Fel, Rémi Cadène, Thomas SerreNeurIPS 2022 · 147 citations
- Interpreting Language Reward Models via Contrastive ExplanationsJunqi Jiang, Tom Bewley, Saumitra Mishra, Freddy Lécué et al.ICLR 2025
