Lune

ICML2026顶会

Generative Online Reinforcement Learning

Chubin Zhang, Zhenglin Wan, Feng Chen, Fuchao Yang, Lang Feng, Yaxin Zhou, Xingrui Yu, Yang You, Ivor Tsang, Bo An

出版方
2026年份
1被引次数

摘要

Reinforcement learning (RL) faces a persistent tension: policies that are stable to optimize (e.g., Gaussians) are often too simple to represent the multimodal action distributions required for complex control. Conversely, expressive generative policies-such as diffusion and flow matchingcan be difficult to optimize in online RL due to intractable likelihoods and gradients propagating through long sampling chains. We address this tension with a key structural principle: decoupling optimization from generation. Building on this, we introduce GORL (Generative Online Reinforcement Learning), an algorithmagnostic framework that trains expressive policies from scratch by confining policy optimization to a tractable latent space while delegating action synthesis to a conditional generative decoder. Using a two-timescale alternating schedule and anchoring decoder refinement to a fixed prior, GORL enables stable optimization while continuously expanding expressiveness. Empirically, GORL consistently outperforms unimodal and generative baselines across diverse continuouscontrol tasks. Notably, GORL achieves returns exceeding 870 on HopperStand, more than 3× the strongest baseline; on high-dimensional humanoid tasks, it further outperforms the strongest non-GORL baseline by over an order of magnitude. Code is available at https://github. com/bennidict23/GoRL.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖