Generative Online Reinforcement Learning
Chubin Zhang, Zhenglin Wan, Feng Chen, Fuchao Yang, Lang Feng, Yaxin Zhou, Xingrui Yu, Yang You, Ivor Tsang, Bo An
摘要
Reinforcement learning (RL) faces a persistent tension: policies that are stable to optimize (e.g., Gaussians) are often too simple to represent the multimodal action distributions required for complex control. Conversely, expressive generative policies-such as diffusion and flow matchingcan be difficult to optimize in online RL due to intractable likelihoods and gradients propagating through long sampling chains. We address this tension with a key structural principle: decoupling optimization from generation. Building on this, we introduce GORL (Generative Online Reinforcement Learning), an algorithmagnostic framework that trains expressive policies from scratch by confining policy optimization to a tractable latent space while delegating action synthesis to a conditional generative decoder. Using a two-timescale alternating schedule and anchoring decoder refinement to a fixed prior, GORL enables stable optimization while continuously expanding expressiveness. Empirically, GORL consistently outperforms unimodal and generative baselines across diverse continuouscontrol tasks. Notably, GORL achieves returns exceeding 870 on HopperStand, more than 3× the strongest baseline; on high-dimensional humanoid tasks, it further outperforms the strongest non-GORL baseline by over an order of magnitude. Code is available at https://github. com/bennidict23/GoRL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
相关 Paper
- Flow-Based Policy for Online Reinforcement LearningLei Lyu, Yunfei Li, Yu Luo, Fuchun Sun 等NeurIPS 2025 · 被引用 39 次
- Behavior Regularization with Flow Latent Policy for Offline Reinforcement LearningYulong Xia, Fuchun SunAAAI 2026
- Flow-Based Single-Step Completion for Efficient and Expressive Policy LearningPrajwal Koirala, Cody FlemingICLR 2026 · 被引用 12 次
- Mean Flow Policy OptimizationXiaoyi Dong, Xi Zhang, Jian ChengICML 2026
- EXPO: Stable Reinforcement Learning with Expressive PoliciesPerry Dong, Qiyang Li, Dorsa Sadigh, Chelsea FinnICLR 2026 · 被引用 35 次
