Breaking the Barrier: Enhanced Utility and Robustness in Smoothed DRL Agents
Chung-En Sun, Sicun Gao, Tsui-Wei Weng
摘要
Robustness remains a paramount concern in deep reinforcement learning (DRL), with randomized smoothing emerging as a key technique for enhancing this attribute. However, a notable gap exists in the performance of current smoothed DRL agents, often characterized by significantly low clean rewards and weak robustness. In response to this challenge, our study introduces innovative algorithms aimed at training effective smoothed robust DRL agents. We propose S-DQN and S-PPO, novel approaches that demonstrate remarkable improvements in clean rewards, empirical robustness, and robustness guarantee across standard RL benchmarks. Notably, our S-DQN and S-PPO agents not only significantly outperform existing smoothed agents by an average factor of under the strongest attack, but also surpass previous robustly-trained agents by an average factor of . This represents a significant leap forward in the field. Furthermore, we introduce Smoothed Attack, which is more effective in decreasing the rewards of smoothed agents than existing adversarial attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Robust Deep Reinforcement Learning against Adversarial Behavior ManipulationShojiro Yamabe, Kazuto Fukuchi, Jun SakumaICLR 2026 · 被引用 1 次
- CAMP in the Odyssey: Provably Robust Reinforcement Learning with Certified Radius MaximizationDerui Wang, Kristen Moore, Diksha Goel, Minjune Kim 等USENIX Security 2025
- Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy EvaluationKosuke Nakanishi, Akihiro Kubo, Yuji Yasui, Shin IshiiICML 2025
- On the Tension Between Optimality and Adversarial Robustness in Policy OptimizationHaoran Li, Jiayu Lv, Congying Han, Zicheng Zhang 等ICLR 2026
它引用的顶会 Paper9
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State ObservationsHuan Zhang, Hongge Chen, Chaowei Xiao, Bo Li 等NeurIPS 2020 · 被引用 437 次
- Robust Reinforcement Learning on State Observations with Learned Optimal AdversaryHuan Zhang, Hongge Chen, Duane S. Boning, Cho-Jui HsiehICLR 2021 · 被引用 212 次
- Denoised Smoothing: A Provable Defense for Pretrained ClassifiersHadi Salman, Mingjie Sun, Greg Yang, Ashish Kapoor 等NeurIPS 2020 · 被引用 191 次
- Robust Deep Reinforcement Learning through Adversarial LossTuomas P. Oikarinen, Wang Zhang, Alexandre Megretski, Luca Daniel 等NeurIPS 2021 · 被引用 134 次
- Who Is the Strongest Enemy? Towards Optimal and Efficient Evasion Attacks in Deep RLYanchao Sun, Ruijie Zheng, Yongyuan Liang, Furong HuangICLR 2022 · 被引用 82 次
相关 Paper
- Belief-Enriched Pessimistic Q-Learning against Adversarial State PerturbationsXiaolin Sun, Zizhan ZhengICLR 2024 · 被引用 4 次
- CROP: Certifying Robust Policies for Reinforcement Learning through Functional SmoothingFan Wu, Linyi Li, Zijian Huang, Yevgeniy Vorobeychik 等ICLR 2022 · 被引用 64 次
- Deep Reinforcement Learning with Robust and Smooth PolicyQianli Shen, Yan Li, Haoming Jiang, Zhaoran Wang 等ICML 2020 · 被引用 95 次
- On the Robustness of Safe Reinforcement Learning under Observational PerturbationsZuxin Liu, Zijian Guo, Zhepeng Cen, Huan Zhang 等ICLR 2023 · 被引用 9 次
- Adversarial Policy Training against Deep Reinforcement LearningXian Wu, Wenbo Guo, Hua Wei, Xinyu XingUSENIX Security 2021 · 被引用 19 次
