Lune

NeurIPS2025顶会

ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning

Tonghe Zhang, Chao Yu, Sichang Su, Yu Wang

2025年份
101被引次数
13顶会引用

摘要

We propose ReinFlow, a simple yet effective online reinforcement learning (RL) framework that fine-tunes a family of flow matching policies for continuous robotic control. Derived from rigorous RL theory, ReinFlow injects learnable noise into a flow policy's deterministic path, converting the flow into a discrete-time Markov Process for exact and straightforward likelihood computation. This conversion facilitates exploration and ensures training stability, enabling ReinFlow to fine-tune diverse flow model variants stably, including Rectified Flow [34] and Shortcut Models [18], particularly at very few or even one denoising step. We benchmark ReinFlow in representative locomotion and manipulation tasks, including longhorizon planning with visual input and sparse reward. The episode reward of Rectified Flow policies obtained an average net growth of 135.36% after fine-tuning in challenging legged locomotion tasks while saving denoising steps and 82.63% of wall time compared to state-of-the-art diffusion RL fine-tuning method DPPO [42]. The success rate of the Shortcut Model policies in state and visual manipulation tasks achieved an average net increase of 40.34% after fine-tuning with ReinFlow at four or even one denoising step, whose performance is comparable to fine-tuned DDIM policies while saving computation time for an average of 23.20% . Code, model, and checkpoints available on the project website: https://reinflow.github.io/ * Corresponding author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

E π +∞ h=0 γ h r h (a h , o h ) . The Q function value function and the advantage function are defined as

(1) We drop the subscript h for V π h , Q π h , and A π h when the policy and POMDP are stationary.

Flow Matching Models Flow matching [32] transforms random variables from one distribution p 0 to another p 1 with flow mappings ψ : [0, 1] × X → X , where X t := ψ t (X 0 ), t ∈ [0, 1]. This process is associated with an ODE:

Practitioners also sample t from the beta distribution [7] or the logit normal distribution [16]. [18] proposes "Shortcut Models" to further improve the generation quality of Rectified Flow at very few denoising steps by enforcing the velocity generated by two steps to align with that generated by a single step. During inference, we numerically solve the transport equation by integrating the learned velocity field:

where K is the number of denoising steps, 0 = t 0 < t 1 < . . . < t K-1 < 1 = t K are the discretized time steps, ∆t i = t i+1 -t i is the step size.

Flow Matching Policy When we instantiate a flow-matching model in the action-generation setting, we obtain a flow-matching policy for robot learning. We denote by a t h the denoised action at time t

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext cf2f66de-0a27-4e1e-9bb1-f44f5ac3a80e

引用它的顶会 Paper13

问问它们各自怎么用它

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖