Lune

ICLR2026Top-tier venue

Offline Preference-Based Value Optimization

Hyungkyu Kang, Min-hwan Oh

2026Year

Abstract

We study the problem of offline preference-based reinforcement learning (PbRL), where the agent learns from pre-collected preference data by comparing trajectory pairs. While prior work has established theoretical foundations for offline PbRL, existing algorithms face significant practical limitations: some rely on computationally intractable optimization procedures, while others suffer from unstable training and high performance variance. To address these challenges, we propose Preference-based Value Optimization (PVO), a simple and practical algorithm that achieves both strong empirical performance and theoretical guarantees. PVO directly optimizes the value function consistent with preference feedback by minimizing a novel value alignment loss. We prove that PVO attains a rate-optimal sample complexity of O(ε−2)\mathcal{O}(\varepsilon^{-2}), and further show that the value alignment loss is applicable not only to value-based methods but also to actor–critic algorithms. Empirically, PVO achieves robust and stable performance across diverse continuous control benchmarks. It consistently outperforms strong baselines, including methods without theoretical guarantees, while requiring no additional hyperparameters for preference learning. Moreover, our ablation study demonstrates that substituting the standard TD loss with the value alignment loss substantially improves learning from preference data, confirming its effectiveness for PbRL.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 31f72b5c-c702-45fa-b84b-3308df888b24

Builds on41

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines