Detecting Adversarial Directions in Deep Reinforcement Learning to Make Robust Decisions
Ezgi Korkmaz, Jonah Brown-Cohen
摘要
Learning in MDPs with highly complex state representations is currently possible due to multiple advancements in reinforcement learning algorithm design. However, this incline in complexity, and furthermore the increase in the dimensions of the observation came at the cost of volatility that can be taken advantage of via adversarial attacks (i.e. moving along worst-case directions in the observation space). To solve this policy instability problem we propose a novel method to detect the presence of these non-robust directions via local quadratic approximation of the deep neural policy loss. Our method provides a theoretical basis for the fundamental cut-off between safe observations and adversarial observations. Furthermore, our technique is computationally efficient, and does not depend on the methods used to produce the worst-case directions. We conduct extensive experiments in the Arcade Learning Environment with several different adversarial attack techniques. Most significantly, we demonstrate the effectiveness of our approach even in the setting where non-robust directions are explicitly optimized to circumvent our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Illusory Attacks: Information-theoretic detectability matters in adversarial attacksTim Franzmeyer, Stephen Marcus McAleer, João F. Henriques, Jakob Nicolaus Foerster 等ICLR 2024 · 被引用 12 次
- Diffusion Guided Adversarial State Perturbations in Reinforcement LearningXiaolin Sun, Feidi Liu, Zhengming Ding, Zizhan ZhengNeurIPS 2025
它引用的顶会 Paper4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
- Stealthy and Efficient Adversarial Attacks against Deep Reinforcement LearningJianwen Sun, Tianwei Zhang, Xiaofei Xie, Lei Ma 等AAAI 2020 · 被引用 141 次
- Deep Reinforcement Learning Policies Learn Shared Adversarial Features across MDPsEzgi KorkmazAAAI 2022 · 被引用 33 次
相关 Paper
- Understanding and Diagnosing Deep Reinforcement LearningEzgi KorkmazICML 2024 · 被引用 10 次
- Adversarial Robust Deep Reinforcement Learning Requires Redefining RobustnessEzgi KorkmazAAAI 2023 · 被引用 38 次
- Efficient Adversarial Training without Attacking: Worst-Case-Aware Robust Reinforcement LearningYongyuan Liang, Yanchao Sun, Ruijie Zheng, Furong HuangNeurIPS 2022 · 被引用 79 次
- How to Lose Inherent Counterfactuality in Reinforcement LearningEzgi KorkmazICLR 2026
- When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPsJose Aguilar Escamilla, Haoyang Hong, Jiawei Li, Haoyu Zhao 等ICML 2026
