Towards Robust and Safe Reinforcement Learning with Benign Off-policy Data
Zuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang, Yihang Yao, Hanjiang Hu, Ding Zhao
摘要
Previous work demonstrates that the optimal safe reinforcement learning policy in a noisefree environment is vulnerable and could be unsafe under observational attacks. While adversarial training effectively improves robustness and safety, collecting samples by attacking the behavior agent online could be expensive or prohibitively dangerous in many applications. We propose the robuSt vAriational ofF-policy lEaRning (SAFER) approach, which only requires benign training data without attacking the agent. SAFER obtains an optimal non-parametric variational policy distribution via convex optimization and then uses it to improve the parameterized policy robustly via supervised learning. The two-stage policy optimization facilitates robust training, and extensive experiments on multiple robot platforms show the efficiency of SAFER in learning a robust and safe policy: achieving the same reward with much fewer constraint violations during training than on-policy baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Constraint-Conditioned Policy Optimization for Versatile Safe Reinforcement LearningYihang Yao, Zuxin Liu, Zhepeng Cen, Jiacheng Zhu 等NeurIPS 2023 · 被引用 24 次
- RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool LearningJunjie Ye, Yilong Wu, Songyang Gao, Caishuang Huang 等EMNLP 2024 · 被引用 6 次
- How Does the Lagrangian Guide Safe Reinforcement Learning through Diffusion Models?Xiaoyuan Cheng, Wenxuan Yuan, Boyang Li, Yuanchao Xu 等ICML 2026
它引用的顶会 Paper14
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State ObservationsHuan Zhang, Hongge Chen, Chaowei Xiao, Bo Li 等NeurIPS 2020 · 被引用 437 次
- Adversarial Policies: Attacking Deep Reinforcement LearningAdam Gleave, Michael Dennis, Cody Wild, Neel Kant 等ICLR 2020 · 被引用 415 次
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 被引用 403 次
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 被引用 244 次
相关 Paper
- On the Robustness of Safe Reinforcement Learning under Observational PerturbationsZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang 等ICLR 2023 · 被引用 9 次
- Constrained Variational Policy Optimization for Safe Reinforcement LearningZuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Wei Liu 等ICML 2022 · 被引用 112 次
- Risk-Averse Offline Reinforcement LearningNúria Armengol Urpí, Sebastian Curi, Andreas KrauseICLR 2021 · 被引用 81 次
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 被引用 127 次
- First Order Constrained Optimization in Policy SpaceYiming Zhang, Quan Vuong, Keith W. RossNeurIPS 2020 · 被引用 238 次
