Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning
Akanksha Atrey, Kaleigh Clary, David D. Jensen
摘要
Saliency maps are frequently used to support explanations of the behavior of deep reinforcement learning (RL) agents. However, a review of how saliency maps are used in practice indicates that the derived explanations are often unfalsifiable and can be highly subjective. We introduce an empirical approach grounded in counterfactual reasoning to test the hypotheses generated from saliency maps and assess the degree to which they correspond to the semantics of RL environments. We use Atari games, a common benchmark for deep RL, to evaluate three types of saliency maps. Our results show the extent to which existing claims about Atari games can be evaluated and suggest that saliency maps are best viewed as an exploratory tool rather than an explanatory tool.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
- When Explanations Lie: Why Many Modified BP Attributions FailLeon Sixt, Maximilian Granz, Tim LandgrafICML 2020 · 被引用 147 次
- Additive MIL: Intrinsically Interpretable Multiple Instance Learning for PathologySyed Ashar Javed, Dinkar Juyal, Harshith Padigela, Amaro Taylor-Weiner 等NeurIPS 2022 · 被引用 124 次
- EDGE: Explaining Deep Reinforcement Learning PoliciesWenbo Guo, Xian Wu, Usmann Khan, Xinyu XingNeurIPS 2021 · 被引用 79 次
- Look where you look! Saliency-guided Q-networks for generalization in visual Reinforcement LearningDavid Bertoin, Adil Zouitine, Mehdi Zouitine, Emmanuel RachelsonNeurIPS 2022 · 被引用 67 次
相关 Paper
- Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature AttributionNikaash Puri, Sukriti Verma, Piyush Gupta, Dhruv Kayastha 等ICLR 2020 · 被引用 99 次
- Machine versus Human Attention in Deep Reinforcement Learning TasksSihang Guo, Ruohan Zhang, Bo Liu, Yifeng Zhu 等NeurIPS 2021 · 被引用 38 次
- Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language ModelsMarvin Pafla, Kate Larson, Mark HancockCHI 2024 · 被引用 16 次
- Counterfactual-based Saliency Map: Towards Visual Contrastive Explanations for Neural NetworksXue Wang, Zhibo Wang, Haiqin Weng, Hengchang Guo 等ICCV 2023 · 被引用 15 次
- Explaining RL Decisions with TrajectoriesShripad Vilasrao Deshmukh, Arpan Dasgupta, Balaji Krishnamurthy, Nan Jiang 等ICLR 2023
