Baffle: Hiding Backdoors in Offline Reinforcement Learning Datasets
Chen Gong, Zhou Yang, Yunpeng Bai, Junda He, Jieke Shi, Kecen Li, Arunesh Sinha, Bowen Xu, Xinwen Hou, David Lo, Tianhao Wang
摘要
Reinforcement learning (RL) makes an agent learn from trial-and-error experiences gathered during the interaction with the environment. Recently, offline RL has become a popular RL paradigm because it saves the interactions with environments. In offline RL, data providers share large pre-collected datasets, and others can train high-quality agents without interacting with the environments. This paradigm has demonstrated effectiveness in critical tasks like robot control, autonomous driving, etc. However, less attention is paid to investigating the security threats to the offline RL system. This paper focuses on backdoor attacks, where some perturbations are added to the data (observations) such that given normal observations, the agent takes high-rewards actions, and low-reward actions on observations injected with triggers. In this paper, we propose Baffle (Backdoor Attack for Offline Reinforcement Learning), an approach that automatically implants backdoors to RL agents by poisoning the offline RL dataset, and evaluate how different offline RL algorithms react to this attack. Our experiments conducted on four tasks and nine offline RL algorithms expose a disquieting fact: none of the existing offline RL algorithms has been immune to such a backdoor attack. More specifically, Baffle modifies 10% of the datasets for four tasks (3 robotic controls and 1 autonomous driving). Agents trained on the poisoned datasets perform well in normal settings. However, when triggers are presented, the agents’ performance decreases drastically by 63.2%, 53.9%, 64.7%, and 47.4% in the four tasks on average. The backdoor still persists after fine-tuning poisoned agents on clean datasets. We further show that the inserted backdoor is also hard to be detected by a popular defensive method. This paper calls attention to developing more effective protection for the open-source offline RL dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based AgentsWenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen 等NeurIPS 2024 · 被引用 195 次
- Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous DrivingXuan Chen, Shiwei Feng, Zikang Xiong, Shengwei An 等NeurIPS 2025 · 被引用 6 次
- Fox in the Henhouse: Supply-Chain Backdoor Attacks Against Reinforcement LearningShijie Liu, Andrew C. Cullen, Paul MONTAGUE, Sarah Erfani 等ICML 2026 · 被引用 5 次
- TrojanTO: Action-Level Backdoor Attacks Against Trajectory Optimization ModelsYang Dai, Oubo Ma, Xingxing Liang, Longfei Zhang 等ICLR 2026 · 被引用 3 次
- Watermarking LLM Agent TrajectoriesWenlong Meng, Chen GONG, Terry Yue Zhuo, Fan Zhang 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper31
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- GraphCodeBERT: Pre-training Code Representations with Data FlowDaya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng 等ICLR 2021 · 被引用 1,644 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- Machine UnlearningLucas Bourtoule, Varun Chandrasekaran, Christopher A. Choquette-Choo, Hengrui Jia 等S&P 2021 · 被引用 1,381 次
相关 Paper
- SleeperNets: Universal Backdoor Poisoning Attacks Against Reinforcement Learning AgentsEthan Rathbun, Christopher Amato, Alina OpreaNeurIPS 2024 · 被引用 27 次
- BadRL: Sparse Targeted Backdoor Attack against Reinforcement LearningJing Cui, Yufei Han, Yuzhe Ma, Jianbin Jiao 等AAAI 2024 · 被引用 31 次
- SHINE: Shielding Backdoors in Deep Reinforcement LearningZhuowen Yuan, Wenbo Guo, Jinyuan Jia, Bo Li 等ICML 2024 · 被引用 4 次
- Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement LearningEthan Rathbun, Wo Wei Lin, Alina Oprea, Christopher AmatoICLR 2026 · 被引用 5 次
- TrojDRL: Evaluation of Backdoor Attacks on Deep Reinforcement LearningPanagiota Kiourti, Kacper Wardega, Susmit Jha, Wenchao LiDAC 2020 · 被引用 72 次
