WCSAC: Worst-Case Soft Actor Critic for Safety-Constrained Reinforcement Learning
Qisong Yang, Thiago D. Simão, Simon H. Tindemans, Matthijs T. J. Spaan
摘要
Safe exploration is regarded as a key priority area for reinforcement learning research. With separate reward and safety signals, it is natural to cast it as constrained reinforcement learning, where expected long-term costs of policies are constrained. However, it can be hazardous to set constraints on the expected safety signal without considering the tail of the distribution. For instance, in safety-critical domains, worst-case analysis is required to avoid disastrous results. We present a novel reinforcement learning algorithm called Worst-Case Soft Actor Critic, which extends the Soft Actor Critic algorithm with a safety critic to achieve risk control. More specifically, a certain level of conditional Value-at-Risk from the distribution is regarded as a safety measure to judge the constraint satisfaction, which guides the change of adaptive safety weights to achieve a trade-off between reward and safety. As a result, we can optimize policies under the premise that their worst-case performance satisfies the constraints. The empirical analysis shows that our algorithm attains better risk control compared to expectation-based methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess 等ICLR 2022 · 被引用 84 次
- Constrained Decision Transformer for Offline Safe Reinforcement LearningZuxin Liu, Zijian Guo, Yihang Yao, Zhepeng Cen 等ICML 2023 · 被引用 82 次
- SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential ModelingYixian Zhang, Shu'ang Yu, Tonghe Zhang, Mo Guang 等ICLR 2026 · 被引用 33 次
- A Provably-Efficient Model-Free Algorithm for Infinite-Horizon Average-Reward Constrained Markov Decision ProcessesHonghao Wei, Xin Liu, Lei YingAAAI 2022 · 被引用 31 次
- CEM: Constrained Entropy Maximization for Task-Agnostic Safe ExplorationQisong Yang, Matthijs T. J. SpaanAAAI 2023 · 被引用 24 次
相关 Paper
- Safe Multi-Agent Reinforcement Learning via Distributional Safety Critic and Maximum Entropy OptimizationQiwei Liu, Ye Yuan, Lingyue Zhang, Kaitian Chen 等AAAI 2026
- Reward Penalties on Augmented States for Solving Richly Constrained RL EffectivelyHao Jiang, Tien Mai, Pradeep Varakantham, Huy HoangAAAI 2024 · 被引用 2 次
- Safety Representations for Safer Policy LearningKaustubh Mani, Vincent Mai, Charlie Gauthier, Annie S. Chen 等ICLR 2025
- Extreme Value Policy Optimization for Safe Reinforcement LearningShiqing Gao, Yihang Zhou, Shuai Shao, Haoyu Luo 等ICML 2025
- Learning Policies with Zero or Bounded Constraint Violation for Constrained MDPsTao Liu, Ruida Zhou, Dileep Kalathil, Panganamala R. Kumar 等NeurIPS 2021 · 被引用 110 次
