Diffusion Guided Adversarial State Perturbations in Reinforcement Learning
Xiaolin Sun, Feidi Liu, Zhengming Ding, Zizhan Zheng
Abstract
Reinforcement learning (RL) systems, while achieving remarkable success across various domains, are vulnerable to adversarial attacks. This is especially a concern in vision-based environments where minor manipulations of high-dimensional image inputs can easily mislead the agent's behavior. To this end, various defenses have been proposed recently, with state-of-the-art approaches achieving robust performance even under large state perturbations. However, after closer investigation, we found that the effectiveness of the current defenses is due to a fundamental weakness of the existing l p norm-constrained attacks, which can barely alter the semantics of image input even under a relatively large perturbation budget. In this work, we propose SHIFT, a novel policy-agnostic diffusion-based state perturbation attack to go beyond this limitation. Our attack is able to generate perturbed states that are semantically different from the true states while remaining realistic and history-aligned to avoid detection. Evaluations show that our attack effectively breaks existing defenses, including the most sophisticated ones, significantly outperforming existing attacks while being more perceptually stealthy. The results highlight the vulnerability of RL agents to semantics-aware adversarial perturbations, indicating the importance of developing more robust policies. Our code can be found at this GitHub Repo.
With these two shortcomings in mind, we propose SHIFT (Stealthy History-alIgned diFfusion aTtack), a novel semanticsaware and policy-agnostic attack method that goes beyond the traditional l p norm constraint. Our approach is grounded in precise definitions of realistic states and three attack properties: semantic-altering, historically-aligned, and trajectory-faithful, which provide novel characterizations of semantics-aware and stealthy attacks in sequential decision-making. As these metrics are computationally expensive to evaluate, we provide practical methods to approximate them. Our main contribution is the development of a diffusion-based attack framework that utilizes classifier-free guidance to approximate history-aligned state generation, which is further improved using policy guidance to generate effective, realistic, and stealthy state perturbations.
Using this novel guided diffusion approach, we propose two versions of SHIFT: SHIFT-O perturbs the image input conditioned on the actual history to immediately induce suboptimal actions, while SHIFT-I guides the agent toward an imagined trajectory that is self-consistent, but ultimately leads to poor performance when actions are executed in the real environment. We compare these two methods in Figure 1, with a working example in the Freeway environment given in Figure 6 in Appendix C. SHIFT is policy-agnostic, capable of adapting to multiple victim policies without retraining, and applicable to both value-based and policy-based reinforcement learning methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c7de38a-ddd3-4011-95b7-c6590842407fBuilds on34
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
Related papers
- Transferable and Stealthy Adversarial Attacks on Large Vision-Language ModelsZhewen Yao, Yao Zhu, Shiliang ZhangICLR 2026 · 2 citations
- Belief-Enriched Pessimistic Q-Learning against Adversarial State PerturbationsXiaolin Sun, Zizhan ZhengICLR 2024 · 4 citations
- SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMsShuhan Xu, Siyuan Liang, Hongling Zheng, Aishan Liu et al.AAAI 2026 · 5 citations
- Who Is the Strongest Enemy? Towards Optimal and Efficient Evasion Attacks in Deep RLYanchao Sun, Ruijie Zheng, Yongyuan Liang, Furong HuangICLR 2022 · 82 citations
- Natural Black-Box Adversarial Examples against Deep Reinforcement LearningMengran Yu, Shiliang SunAAAI 2022 · 15 citations
