Finding and Visualizing Weaknesses of Deep Reinforcement Learning Agents
Christian Rupprecht, Cyril Ibrahim, Christopher J. Pal
摘要
As deep reinforcement learning driven by visual perception becomes more widely used there is a growing need to better understand and probe the learned agents. Understanding the decision making process and its relationship to visual inputs can be very valuable to identify problems in learned behavior. However, this topic has been relatively under-explored in the research community. In this work we present a method for synthesizing visual inputs of interest for a trained agent. Such inputs or states could be situations in which specific actions are necessary. Further, critical states in which a very high or a very low reward can be achieved are often interesting to understand the situational awareness of the system as they can correspond to risky states. To this end, we learn a generative model over the state space of the environment and use its latent space to optimize a target function for the state of interest. In our experiments we show that this method can generate insights for a variety of environments and reinforcement learning methods. We explore results in the standard Atari benchmark games as well as in an autonomous driving simulator. Based on the efficiency with which we have been able to identify behavioural weaknesses with this technique, we believe this general approach could serve as an important tool for AI safety applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Accountability in Offline Reinforcement Learning: Explaining Decisions with a Corpus of ExamplesHao Sun, Alihan Hüyük, Daniel Jarrett, Mihaela van der SchaarNeurIPS 2023 · 被引用 13 次
- On the Safety of Interpretable Machine Learning: A Maximum Deviation ApproachDennis Wei, Rahul Nair, Amit Dhurandhar, Kush R. Varshney 等NeurIPS 2022 · 被引用 12 次
- Understanding the Evolution of Linear Regions in Deep Reinforcement LearningSetareh Cohan, Nam Hee Kim, David Rolnick, Michiel van de PanneNeurIPS 2022 · 被引用 10 次
- Local Explanations for Reinforcement LearningRonny Luss, Amit Dhurandhar, Miao LiuAAAI 2023 · 被引用 5 次
相关 Paper
- Learning Human Objectives by Evaluating Hypothetical BehaviorSiddharth Reddy, Anca D. Dragan, Sergey Levine, Shane Legg 等ICML 2020 · 被引用 81 次
- Machine versus Human Attention in Deep Reinforcement Learning TasksSihang Guo, Ruohan Zhang, Bo Liu, Yifeng Zhu 等NeurIPS 2021 · 被引用 38 次
- Test Where Decisions Matter: Importance-driven Testing for Deep Reinforcement LearningStefan Pranger, Hana Chockler, Martin Tappler, Bettina KönighoferNeurIPS 2024 · 被引用 7 次
- Safe Reinforcement Learning From Pixels Using a Stochastic Latent RepresentationYannick Hogewind, Thiago D. Simão, Tal Kachman, Nils JansenICLR 2023 · 被引用 3 次
- DeepSynth: Automata Synthesis for Automatic Task Segmentation in Deep Reinforcement LearningMohammadhosein Hasanbeig, Natasha Yogananda Jeppu, Alessandro Abate, Tom Melham 等AAAI 2021 · 被引用 62 次
