Adaptive Reward-Poisoning Attacks against Reinforcement Learning
Xuezhou Zhang, Yuzhe Ma, Adish Singla, Xiaojin Zhu
Abstract
In reward-poisoning attacks against reinforcement learning (RL), an attacker can perturb the environment reward into at each step, with the goal of forcing the RL agent to learn a nefarious policy. We categorize such attacks by the infinity-norm constraint on : We provide a lower threshold below which reward-poisoning attack is infeasible and RL is certified to be safe; we provide a corresponding upper threshold above which the attack is feasible. Feasible attacks can be further categorized as non-adaptive where depends only on , or adaptive where depends further on the RL agent's learning process at time . Non-adaptive attacks have been the focus of prior works. However, we show that under mild conditions, adaptive attacks can achieve the nefarious policy in steps polynomial in state-space size , whereas non-adaptive attacks require exponential steps. We provide a constructive proof that a Fast Adaptive Attack strategy achieves the polynomial rate. Finally, we show that empirically an attacker can find effective reward-poisoning attacks using state-of-the-art deep RL techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 89fbbb07-2916-49a2-b1ab-486e887f959cCited by top-tier papers34
- Adversarial Attacks on Adversarial BanditsYuzhe Ma, Zhijin ZhouICLR 2023 · 199 citations
- Efficient Adversarial Training without Attacking: Worst-Case-Aware Robust Reinforcement LearningYongyuan Liang, Yanchao Sun, Ruijie Zheng, Furong HuangNeurIPS 2022 · 79 citations
- Vulnerability-Aware Poisoning Mechanism for Online RL with Unknown DynamicsYanchao Sun, Da Huo, Furong HuangICLR 2021 · 57 citations
- Learning to Attack Federated Learning: A Model-based Reinforcement Learning Attack FrameworkHenger Li, Xiaolin Sun, Zizhan ZhengNeurIPS 2022 · 55 citations
- Provably Efficient Black-Box Action Poisoning Attacks Against Reinforcement LearningGuanlin Liu, Lifeng LaiNeurIPS 2021 · 55 citations
Related papers
- When Can You Poison Rewards? A Tight Characterization of Reward Poisoning in Linear MDPsJose Aguilar Escamilla, Haoyang Hong, Jiawei Li, Haoyu Zhao et al.ICML 2026
- Policy Teaching via Environment Poisoning: Training-time Adversarial Attacks against Reinforcement LearningAmin Rakhsha, Goran Radanovic, Rati Devidze, Xiaojin Zhu et al.ICML 2020 · 145 citations
- On the Robustness of Safe Reinforcement Learning under Observational PerturbationsZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang et al.ICLR 2023 · 9 citations
- Efficient Adversarial Attacks on Online Multi-agent Reinforcement LearningGuanlin Liu, Lifeng LaiNeurIPS 2023 · 24 citations
- Adversarial Inception Backdoor Attacks against Reinforcement LearningEthan Rathbun, Alina Oprea, Christopher AmatoICML 2025
