Diffusion Guided Adversarial State Perturbations in Reinforcement Learning
Xiaolin Sun, Feidi Liu, Zhengming Ding, Zizhan Zheng
摘要
Reinforcement learning (RL) systems, while achieving remarkable success across various domains, are vulnerable to adversarial attacks. This is especially a concern in vision-based environments where minor manipulations of high-dimensional image inputs can easily mislead the agent's behavior. To this end, various defenses have been proposed recently, with state-of-the-art approaches achieving robust performance even under large state perturbations. However, after closer investigation, we found that the effectiveness of the current defenses is due to a fundamental weakness of the existing l p norm-constrained attacks, which can barely alter the semantics of image input even under a relatively large perturbation budget. In this work, we propose SHIFT, a novel policy-agnostic diffusion-based state perturbation attack to go beyond this limitation. Our attack is able to generate perturbed states that are semantically different from the true states while remaining realistic and history-aligned to avoid detection. Evaluations show that our attack effectively breaks existing defenses, including the most sophisticated ones, significantly outperforming existing attacks while being more perceptually stealthy. The results highlight the vulnerability of RL agents to semantics-aware adversarial perturbations, indicating the importance of developing more robust policies. Our code can be found at this GitHub Repo.
With these two shortcomings in mind, we propose SHIFT (Stealthy History-alIgned diFfusion aTtack), a novel semanticsaware and policy-agnostic attack method that goes beyond the traditional l p norm constraint. Our approach is grounded in precise definitions of realistic states and three attack properties: semantic-altering, historically-aligned, and trajectory-faithful, which provide novel characterizations of semantics-aware and stealthy attacks in sequential decision-making. As these metrics are computationally expensive to evaluate, we provide practical methods to approximate them. Our main contribution is the development of a diffusion-based attack framework that utilizes classifier-free guidance to approximate history-aligned state generation, which is further improved using policy guidance to generate effective, realistic, and stealthy state perturbations.
Using this novel guided diffusion approach, we propose two versions of SHIFT: SHIFT-O perturbs the image input conditioned on the actual history to immediately induce suboptimal actions, while SHIFT-I guides the agent toward an imagined trajectory that is self-consistent, but ultimately leads to poor performance when actions are executed in the real environment. We compare these two methods in Figure 1, with a working example in the Freeway environment given in Figure 6 in Appendix C. SHIFT is policy-agnostic, capable of adapting to multiple victim policies without retraining, and applicable to both value-based and policy-based reinforcement learning methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper34
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 被引用 1,115 次
相关 Paper
- Transferable and Stealthy Adversarial Attacks on Large Vision-Language ModelsZhewen Yao, Yao Zhu, Shiliang ZhangICLR 2026 · 被引用 2 次
- Belief-Enriched Pessimistic Q-Learning against Adversarial State PerturbationsXiaolin Sun, Zizhan ZhengICLR 2024 · 被引用 4 次
- SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMsShuhan Xu, Siyuan Liang, Hongling Zheng, Aishan Liu 等AAAI 2026 · 被引用 5 次
- Who Is the Strongest Enemy? Towards Optimal and Efficient Evasion Attacks in Deep RLYanchao Sun, Ruijie Zheng, Yongyuan Liang, Furong HuangICLR 2022 · 被引用 82 次
- Natural Black-Box Adversarial Examples against Deep Reinforcement LearningMengran Yu, Shiliang SunAAAI 2022 · 被引用 15 次
