State2Explanation: Concept-Based Explanations to Benefit Agent Learning and User Understanding
Devleena Das, Sonia Chernova, Been Kim
Abstract
As more non-AI experts use complex AI systems for daily tasks, there has been an increasing effort to develop methods that produce explanations of AI decision making that are understandable by non-AI experts. Towards this effort, leveraging higher-level concepts and producing concept-based explanations have become a popular method. Most concept-based explanations have been developed for classification techniques, and we posit that the few existing methods for sequential decision making are limited in scope. In this work, we first contribute a desiderata for defining "concepts" in sequential decision making settings. Additionally, inspired by the Protégé Effect which states explaining knowledge often reinforces one's self-learning, we explore how concept-based explanations of an RL agent's decision making can in turn improve the agent's learning rate, as well as improve end-user understanding of the agent's decision making. To this end, we contribute a unified framework, State2Explanation (S2E), that involves learning a joint embedding model between state-action pairs and concept-based explanations, and leveraging such learned model to both (1) inform reward shaping during an agent's training, and (2) provide explanations to end-users at deployment for improved task performance. Our experimental validations, in Connect 4 and Lunar Lander, demonstrate the success of S2E in providing a dual-benefit, successfully informing reward shaping and improving agent learning rate, as well as significantly improving end user task performance at deployment time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4dc3c8bf-1379-4883-bb0a-4afb22fea1b6Cited by top-tier papers5
- Auxiliary Losses for Learning Generalizable Concept-based ModelsIvaxi Sheth, Samira Ebrahimi KahouNeurIPS 2023 · 52 citations
- Autonomous Capability Assessment of Sequential Decision-Making Systems in Stochastic SettingsPulkit Verma, Rushang Karia, Siddharth SrivastavaNeurIPS 2023 · 14 citations
- Agua: A Concept-Based Explainer for Learning-Enabled SystemsSagar Patel, Dongsu Han, Nina Narodytska, Sangeetha Abdu JyothiSIGCOMM 2025 · 2 citations
- Flexible Concept Bottleneck ModelXingbo Du, Qiantong Dou, Lei Fan, Rui ZhangAAAI 2026
- LICORICE: Label-Efficient Concept-Based Interpretable Reinforcement LearningZhuorui Ye, Stephanie Milani, Geoffrey J. Gordon, Fei FangICLR 2025
Builds on7
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 408 citations
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang et al.NeurIPS 2020 · 256 citations
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 115 citations
- Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable RepresentationsSarath Sreedharan, Utkarsh Soni, Mudit Verma, Siddharth Srivastava et al.ICLR 2022 · 39 citations
Related papers
- Explaining Reinforcement Learning with Shapley ValuesDaniel Beechey, Thomas M. S. Smith, Özgür SimsekICML 2023 · 41 citations
- Teaching Humans When to Defer to a Classifier via ExemplarsHussein Mozannar, Arvind Satyanarayan, David A. SontagAAAI 2022 · 49 citations
- Translate Policy to Language: Flow Matching Generated Rewards for LLM ExplanationsXinyi Yang, Liang Zeng, Heng Dong, Chao Yu et al.ICLR 2026 · 6 citations
- This State Looks Like That: Self-Interpretable Reinforcement Learning Agents using Prototype Soft Actor-CriticAndrea Marzo, Alessio Ragno, Roberto CapobiancoICML 2026
- Learning to Select Prototypical Parts for Interpretable Sequential Data ModelingYifei Zhang, Neng Gao, Cunqing MaAAAI 2023 · 10 citations
