Robust Deep Reinforcement Learning through Bootstrapped Opportunistic Curriculum
Junlin Wu, Yevgeniy Vorobeychik
Abstract
Despite considerable advances in deep reinforcement learning, it has been shown to be highly vulnerable to adversarial perturbations to state observations. Recent efforts that have attempted to improve adversarial robustness of reinforcement learning can nevertheless tolerate only very small perturbations, and remain fragile as perturbation size increases. We propose Bootstrapped Opportunistic Adversarial Curriculum Learning (BCL), a novel flexible adversarial curriculum learning framework for robust reinforcement learning. Our framework combines two ideas: conservatively bootstrapping each curriculum phase with highest quality solutions obtained from multiple runs of the previous phase, and opportunistically skipping forward in the curriculum. In our experiments we show that the proposed BCL framework enables dramatic improvements in robustness of learned policies to adversarial perturbations. The greatest improvement is for Pong, where our framework yields robustness to perturbations of up to 25/255; in contrast, the best existing approach can only tolerate adversarial noise up to 5/255. Our code is available at: https://github.com/jlwu002/BCL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f033d793-ef54-4dbc-b7e2-bd0639cf2b6cCited by top-tier papers8
- BIRD: Generalizable Backdoor Detection and Removal for Deep Reinforcement LearningXuan Chen, Wenbo Guo, Guanhong Tao, Xiangyu Zhang et al.NeurIPS 2023 · 15 citations
- Deep Multitask Learning with Progressive Parameter SharingHaosen Shi, Shen Ren, Tianwei Zhang, Sinno Jialin PanICCV 2023 · 15 citations
- Verified Safe Reinforcement Learning for Neural Network Dynamic ModelsJunlin Wu, Huan Zhang, Yevgeniy VorobeychikNeurIPS 2024 · 13 citations
- Improve Robustness of Reinforcement Learning against Observation Perturbations via l∞ Lipschitz Policy NetworksBuqing Nie, Jingtian Ji, Yangqing Fu, Yue GaoAAAI 2024 · 10 citations
- Bootstrapped Policy Learning for Task-oriented Dialogue through Goal ShapingYangyang Zhao, Ben Niu, Mehdi Dastani, Shihan WangEMNLP 2024
Builds on11
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State ObservationsHuan Zhang, Hongge Chen, Chaowei Xiao, Bo Li et al.NeurIPS 2020 · 437 citations
- Robust Reinforcement Learning on State Observations with Learned Optimal AdversaryHuan Zhang, Hongge Chen, Duane S. Boning, Cho-Jui HsiehICLR 2021 · 212 citations
Related papers
- Curriculum Reinforcement Learning via Constrained Optimal TransportPascal Klink, Haoyi Yang, Carlo D'Eramo, Jan Peters et al.ICML 2022 · 44 citations
- Efficient Adversarial Training without Attacking: Worst-Case-Aware Robust Reinforcement LearningYongyuan Liang, Yanchao Sun, Ruijie Zheng, Furong HuangNeurIPS 2022 · 79 citations
- On the Robustness of Safe Reinforcement Learning under Observational PerturbationsZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang et al.ICLR 2023 · 9 citations
- Robust Policy Gradient against Strong Data CorruptionXuezhou Zhang, Yiding Chen, Xiaojin Zhu, Wen SunICML 2021 · 43 citations
- Adaptive Procedural Task Generation for Hard-Exploration ProblemsKuan Fang, Yuke Zhu, Silvio Savarese, Li Fei-FeiICLR 2021 · 36 citations
