State2Explanation: Concept-Based Explanations to Benefit Agent Learning and User Understanding
Devleena Das, Sonia Chernova, Been Kim
摘要
As more non-AI experts use complex AI systems for daily tasks, there has been an increasing effort to develop methods that produce explanations of AI decision making that are understandable by non-AI experts. Towards this effort, leveraging higher-level concepts and producing concept-based explanations have become a popular method. Most concept-based explanations have been developed for classification techniques, and we posit that the few existing methods for sequential decision making are limited in scope. In this work, we first contribute a desiderata for defining "concepts" in sequential decision making settings. Additionally, inspired by the Protégé Effect which states explaining knowledge often reinforces one's self-learning, we explore how concept-based explanations of an RL agent's decision making can in turn improve the agent's learning rate, as well as improve end-user understanding of the agent's decision making. To this end, we contribute a unified framework, State2Explanation (S2E), that involves learning a joint embedding model between state-action pairs and concept-based explanations, and leveraging such learned model to both (1) inform reward shaping during an agent's training, and (2) provide explanations to end-users at deployment for improved task performance. Our experimental validations, in Connect 4 and Lunar Lander, demonstrate the success of S2E in providing a dual-benefit, successfully informing reward shaping and improving agent learning rate, as well as significantly improving end user task performance at deployment time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Auxiliary Losses for Learning Generalizable Concept-based ModelsIvaxi Sheth, Samira Ebrahimi KahouNeurIPS 2023 · 被引用 52 次
- Autonomous Capability Assessment of Sequential Decision-Making Systems in Stochastic SettingsPulkit Verma, Rushang Karia, Siddharth SrivastavaNeurIPS 2023 · 被引用 14 次
- Agua: A Concept-Based Explainer for Learning-Enabled SystemsSagar Patel, Dongsu Han, Nina Narodytska, Sangeetha Abdu JyothiSIGCOMM 2025 · 被引用 2 次
- Flexible Concept Bottleneck ModelXingbo Du, Qiantong Dou, Lei Fan, Rui ZhangAAAI 2026
- LICORICE: Label-Efficient Concept-Based Interpretable Reinforcement LearningZhuorui Ye, Stephanie Milani, Geoffrey J. Gordon, Fei FangICLR 2025
它引用的顶会 Paper7
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Explainable Reinforcement Learning through a Causal LensPrashan Madumal, Tim Miller, Liz Sonenberg, Frank VetereAAAI 2020 · 被引用 408 次
- Learning to Utilize Shaping Rewards: A New Approach of Reward ShapingYujing Hu, Weixun Wang, Hangtian Jia, Yixiang Wang 等NeurIPS 2020 · 被引用 256 次
- Towards Unsupervised Image Captioning With Shared Multimodal EmbeddingsIro Laina, Christian Rupprecht, Nassir NavabICCV 2019 · 被引用 115 次
- Bridging the Gap: Providing Post-Hoc Symbolic Explanations for Sequential Decision-Making Problems with Inscrutable RepresentationsSarath Sreedharan, Utkarsh Soni, Mudit Verma, Siddharth Srivastava 等ICLR 2022 · 被引用 39 次
相关 Paper
- Explaining Reinforcement Learning with Shapley ValuesDaniel Beechey, Thomas M. S. Smith, Özgür SimsekICML 2023 · 被引用 41 次
- Teaching Humans When to Defer to a Classifier via ExemplarsHussein Mozannar, Arvind Satyanarayan, David A. SontagAAAI 2022 · 被引用 49 次
- Translate Policy to Language: Flow Matching Generated Rewards for LLM ExplanationsXinyi Yang, Liang Zeng, Heng Dong, Chao Yu 等ICLR 2026 · 被引用 6 次
- This State Looks Like That: Self-Interpretable Reinforcement Learning Agents using Prototype Soft Actor-CriticAndrea Marzo, Alessio Ragno, Roberto CapobiancoICML 2026
- Learning to Select Prototypical Parts for Interpretable Sequential Data ModelingYifei Zhang, Neng Gao, Cunqing MaAAAI 2023 · 被引用 10 次
