Generative Online Reinforcement Learning
Chubin Zhang, Zhenglin Wan, Feng Chen, Fuchao Yang, Lang Feng, Yaxin Zhou, Xingrui Yu, Yang You, Ivor Tsang, Bo An
Abstract
Reinforcement learning (RL) faces a persistent tension: policies that are stable to optimize (e.g., Gaussians) are often too simple to represent the multimodal action distributions required for complex control. Conversely, expressive generative policies-such as diffusion and flow matchingcan be difficult to optimize in online RL due to intractable likelihoods and gradients propagating through long sampling chains. We address this tension with a key structural principle: decoupling optimization from generation. Building on this, we introduce GORL (Generative Online Reinforcement Learning), an algorithmagnostic framework that trains expressive policies from scratch by confining policy optimization to a tractable latent space while delegating action synthesis to a conditional generative decoder. Using a two-timescale alternating schedule and anchoring decoder refinement to a fixed prior, GORL enables stable optimization while continuously expanding expressiveness. Empirically, GORL consistently outperforms unimodal and generative baselines across diverse continuouscontrol tasks. Notably, GORL achieves returns exceeding 870 on HopperStand, more than 3× the strongest baseline; on high-dimensional humanoid tasks, it further outperforms the strongest non-GORL baseline by over an order of magnitude. Code is available at https://github. com/bennidict23/GoRL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f91aeeb-3d70-4948-acb2-cc6b78c06800Builds on17
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Flow-Based Policy for Online Reinforcement LearningLei Lyu, Yunfei Li, Yu Luo, Fuchun Sun et al.NeurIPS 2025 · 39 citations
- Behavior Regularization with Flow Latent Policy for Offline Reinforcement LearningYulong Xia, Fuchun SunAAAI 2026
- Flow-Based Single-Step Completion for Efficient and Expressive Policy LearningPrajwal Koirala, Cody FlemingICLR 2026 · 12 citations
- Mean Flow Policy OptimizationXiaoyi Dong, Xi Zhang, Jian ChengICML 2026
- EXPO: Stable Reinforcement Learning with Expressive PoliciesPerry Dong, Qiyang Li, Dorsa Sadigh, Chelsea FinnICLR 2026 · 35 citations
