Uni-RL: Unifying Online and Offline RL via Implicit Value Regularization
Haoran Xu, Liyuan Mao, Hui Jin, Weinan Zhang, Xianyuan Zhan, Amy Zhang
摘要
The practical implementations of reinforcement learning (RL) often face diverse settings, such as online, offline, and offline-to-online learning. Instead of developing separate algorithms for each setting, we propose Uni-RL, a unified model-free RL framework that addresses all these scenarios within a single formulation. Uni-RL builds on the Implicit Value Regularization (IVR) framework (Xu et al., 2023) and generalizes its dataset behavior constraint to the constraint w.r.t. a reference policy, yielding a unified value learning objective for general settings. The reference policy is chosen to be the target policy in the online setting and the behavior policy in the offline setting. Using an iteratively refined behavior policy solves the over-conservative issue of directly applying IVR in the online setting, it provides an implicit trust-region style update through the value function while being off-policy. Uni-RL also introduces a unified policy extraction objective that estimates in-sample policy gradient using only actions from the reference policy. This not only supports various policy classes, but also theoretically guarantees less value estimation error and larger performance improvement over the reference policy. We evaluate Uni-RL on a range of standard RL benchmarks across online, offline, and offline-to-online settings. In online RL, Uni-RL achieves higher sample efficiency than both off-policy methods without trust-region updates and on-policy methods with trust-region updates. In offline RL, Uni-RL retains the benefits of in-sample learning while outperforming IVR through better policy extraction. In offline-to-online RL, Uni-RL beats both constraint-based methods and unconstrained approaches by effectively balancing stability and adaptability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Value Improved Actor Critic AlgorithmsYaniv Oren, Moritz A. Zanger, Pascal R. van der Vaart, Mustafa Mert Çelikok 等NeurIPS 2025 · 被引用 7 次
- Reinforcement Learning via Value Gradient FlowHaoran Xu, Kaiwen Hu, Somayeh Sojoudi, Amy ZhangICLR 2026 · 被引用 4 次
- From Static Constraints to Dynamic Adaptation: Sample-Level Constraint Relaxation for Offline-to-Online Reinforcement LearningLipeng Zu, YU QIAN, Shayok Chakraborty, Xiaonan ZhangICML 2026
它引用的顶会 Paper35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 被引用 1,402 次
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
相关 Paper
- Uni-O4: Unifying Online and Offline Deep Reinforcement Learning with Multi-Step On-Policy OptimizationKun Lei, Zhengmao He, Chenhao Lu, Kaizhe Hu 等ICLR 2024 · 被引用 31 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- A Clean Slate for Offline Reinforcement LearningMatthew Thomas Jackson, Uljad Berdica, Jarek Liesen, Shimon Whiteson 等NeurIPS 2025 · 被引用 11 次
- VIPO: Value Function Inconsistency Penalized Offline Reinforcement LearningXuyang Chen, Keyu Yan, Guojian Wang, Lin ZhaoICML 2026 · 被引用 3 次
- A Unified Framework for Alternating Offline Model Training and Policy LearningShentao Yang, Shujian Zhang, Yihao Feng, Mingyuan ZhouNeurIPS 2022 · 被引用 18 次
