Lune

ICML2026Top-tier venue

Learning Reward–Cost Balance in Safe RL via Score-Based World Models

Yuetian Wang, Dianxi Shi, Yuanze Wang, Huanhuan Yang, Shiming Song, Chunping Qiu

2026Year

Abstract

Safe reinforcement learning (Safe RL) seeks to optimize long-term performance while ensuring adherence to safety constraints. However, most existing approaches address safety in a simplified manner, typically by linearly combining rewards and costs, which provides limited guidance when safety and performance interact in complex, nonlinear ways. We present USB-RL (Unsupervised Score-Balanced Reinforcement Learning), a model-based framework that learns implicit safety-performance preferences directly from experience. Our approach infers a monotone partialorder score through self-supervised pairwise comparisons of long-horizon outcomes-requiring no human preference labels-capturing nuanced trade-offs without relying on manually tuned cost weights. The learned score guides model-based policy optimization by dynamically balancing safety and performance, enabling flexible and adaptive multi-step planning in imagination-based control. Across diverse safety benchmarks, USB-RL achieves strong returns while substantially reducing safety violations, demonstrating stable and interpretable safety-performance trade-offs.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bf86bc0a-d774-4c49-bfb2-f2b83e5ad348

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines