Lune

NeurIPS2025Top-tier venue

ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning

Tonghe Zhang, Chao Yu, Sichang Su, Yu Wang

2025Year
101Citations
13Top-tier citations

Abstract

We propose ReinFlow, a simple yet effective online reinforcement learning (RL) framework that fine-tunes a family of flow matching policies for continuous robotic control. Derived from rigorous RL theory, ReinFlow injects learnable noise into a flow policy's deterministic path, converting the flow into a discrete-time Markov Process for exact and straightforward likelihood computation. This conversion facilitates exploration and ensures training stability, enabling ReinFlow to fine-tune diverse flow model variants stably, including Rectified Flow [34] and Shortcut Models [18], particularly at very few or even one denoising step. We benchmark ReinFlow in representative locomotion and manipulation tasks, including longhorizon planning with visual input and sparse reward. The episode reward of Rectified Flow policies obtained an average net growth of 135.36% after fine-tuning in challenging legged locomotion tasks while saving denoising steps and 82.63% of wall time compared to state-of-the-art diffusion RL fine-tuning method DPPO [42]. The success rate of the Shortcut Model policies in state and visual manipulation tasks achieved an average net increase of 40.34% after fine-tuning with ReinFlow at four or even one denoising step, whose performance is comparable to fine-tuned DDIM policies while saving computation time for an average of 23.20% . Code, model, and checkpoints available on the project website: https://reinflow.github.io/ * Corresponding author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

E π +∞ h=0 γ h r h (a h , o h ) . The Q function value function and the advantage function are defined as

(1) We drop the subscript h for V π h , Q π h , and A π h when the policy and POMDP are stationary.

Flow Matching Models Flow matching [32] transforms random variables from one distribution p 0 to another p 1 with flow mappings ψ : [0, 1] × X → X , where X t := ψ t (X 0 ), t ∈ [0, 1]. This process is associated with an ODE:

Practitioners also sample t from the beta distribution [7] or the logit normal distribution [16]. [18] proposes "Shortcut Models" to further improve the generation quality of Rectified Flow at very few denoising steps by enforcing the velocity generated by two steps to align with that generated by a single step. During inference, we numerically solve the transport equation by integrating the learned velocity field:

where K is the number of denoising steps, 0 = t 0 < t 1 < . . . < t K-1 < 1 = t K are the discretized time steps, ∆t i = t i+1 -t i is the step size.

Flow Matching Policy When we instantiate a flow-matching model in the action-generation setting, we obtain a flow-matching policy for robot learning. We denote by a t h the denoised action at time t

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext cf2f66de-0a27-4e1e-9bb1-f44f5ac3a80e

Cited by top-tier papers13

Ask how each one uses it

Builds on17

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines