Learning Reward–Cost Balance in Safe RL via Score-Based World Models
Yuetian Wang, Dianxi Shi, Yuanze Wang, Huanhuan Yang, Shiming Song, Chunping Qiu
Abstract
Safe reinforcement learning (Safe RL) seeks to optimize long-term performance while ensuring adherence to safety constraints. However, most existing approaches address safety in a simplified manner, typically by linearly combining rewards and costs, which provides limited guidance when safety and performance interact in complex, nonlinear ways. We present USB-RL (Unsupervised Score-Balanced Reinforcement Learning), a model-based framework that learns implicit safety-performance preferences directly from experience. Our approach infers a monotone partialorder score through self-supervised pairwise comparisons of long-horizon outcomes-requiring no human preference labels-capturing nuanced trade-offs without relying on manually tuned cost weights. The learned score guides model-based policy optimization by dynamically balancing safety and performance, enabling flexible and adaptive multi-step planning in imagination-based control. Across diverse safety benchmarks, USB-RL achieves strong returns while substantially reducing safety violations, demonstrating stable and interpretable safety-performance trade-offs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf86bc0a-d774-4c49-bfb2-f2b83e5ad348Builds on11
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Responsive Safety in Reinforcement Learning by PID Lagrangian MethodsAdam Stooke, Joshua Achiam, Pieter AbbeelICML 2020 · 403 citations
- Temporal Difference Learning for Model Predictive ControlNicklas Hansen, Hao Su, Xiaolong WangICML 2022 · 388 citations
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
Related papers
- Enhancing Efficiency of Safe Reinforcement Learning via Sample ManipulationShangding Gu, Laixi Shi, Yuhao Ding, Alois Knoll et al.NeurIPS 2024 · 14 citations
- Implicit Safety Alignment from Crowd PreferencesQian Lin, Daniel S BrownICML 2026
- An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement LearningQian Lin, Zongkai Liu, Danying Mo, Chao YuNeurIPS 2024 · 8 citations
- Safe Reinforcement Learning by Imagining the Near FutureGarrett Thomas, Yuping Luo, Tengyu MaNeurIPS 2021 · 118 citations
- Towards Safe Reinforcement Learning with a Safety Editor PolicyHaonan Yu, Wei Xu, Haichao ZhangNeurIPS 2022 · 50 citations
