Deep Reinforcement Learning Policies Learn Shared Adversarial Features across MDPs
Ezgi Korkmaz
摘要
The use of deep neural networks as function approximators has led to striking progress for reinforcement learning algorithms and applications. Yet the knowledge we have on decision boundary geometry and the loss landscape of neural policies is still quite limited. In this paper, we propose a framework to investigate the decision boundary and loss landscape similarities across states and across MDPs. We conduct experiments in various games from Arcade Learning Environment, and discover that high sensitivity directions for neural policies are correlated across MDPs. We argue that these high sensitivity directions support the hypothesis that non-robust features are shared across training environments of reinforcement learning agents. We believe our results reveal fundamental properties of the environments used in deep reinforcement learning training, and represent a tangible step towards building robust and reliable deep reinforcement learning agents.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Detecting Adversarial Directions in Deep Reinforcement Learning to Make Robust DecisionsEzgi Korkmaz, Jonah Brown-CohenICML 2023 · 被引用 16 次
- Understanding and Diagnosing Deep Reinforcement LearningEzgi KorkmazICML 2024 · 被引用 10 次
- On the Robustness of Safe Reinforcement Learning under Observational PerturbationsZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang 等ICLR 2023 · 被引用 9 次
- Fuz-RL: A Fuzzy-Guided Robust Framework for Safe Reinforcement Learning under UncertaintyXu Wan, Chao Yang, Cheng Yang, Jie Song 等NeurIPS 2025 · 被引用 1 次
- SEBA: Sample-Efficient Black-Box Attacks on Visual Reinforcement LearningTairan Huang, Yulin Jin, Junxu Liu, Qingqing Ye 等CVPR 2026
它引用的顶会 Paper7
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State ObservationsHuan Zhang, Hongge Chen, Chaowei Xiao, Bo Li 等NeurIPS 2020 · 被引用 437 次
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio 等ICML 2020 · 被引用 303 次
- Stealthy and Efficient Adversarial Attacks against Deep Reinforcement LearningJianwen Sun, Tianwei Zhang, Xiaofei Xie, Lei Ma 等AAAI 2020 · 被引用 141 次
相关 Paper
- Understanding Plasticity in Neural NetworksClare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Ávila Pires 等ICML 2023 · 被引用 162 次
- Adversarial Robust Deep Reinforcement Learning Requires Redefining RobustnessEzgi KorkmazAAAI 2023 · 被引用 38 次
- Flat Reward in Policy Parameter Space Implies Robust Reinforcement LearningHyun-Kyu Lee, Sung Whan YoonICLR 2025
- The Geometry of Robust Value FunctionsKaixin Wang, Navdeep Kumar, Kuangqi Zhou, Bryan Hooi 等ICML 2022 · 被引用 7 次
- Can Neural Nets Learn the Same Model Twice? Investigating Reproducibility and Double Descent from the Decision Boundary PerspectiveGowthami Somepalli, Liam Fowl, Arpit Bansal, Ping-Yeh Chiang 等CVPR 2022
