ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
Tonghe Zhang, Chao Yu, Sichang Su, Yu Wang
Abstract
We propose ReinFlow, a simple yet effective online reinforcement learning (RL) framework that fine-tunes a family of flow matching policies for continuous robotic control. Derived from rigorous RL theory, ReinFlow injects learnable noise into a flow policy's deterministic path, converting the flow into a discrete-time Markov Process for exact and straightforward likelihood computation. This conversion facilitates exploration and ensures training stability, enabling ReinFlow to fine-tune diverse flow model variants stably, including Rectified Flow [34] and Shortcut Models [18], particularly at very few or even one denoising step. We benchmark ReinFlow in representative locomotion and manipulation tasks, including longhorizon planning with visual input and sparse reward. The episode reward of Rectified Flow policies obtained an average net growth of 135.36% after fine-tuning in challenging legged locomotion tasks while saving denoising steps and 82.63% of wall time compared to state-of-the-art diffusion RL fine-tuning method DPPO [42]. The success rate of the Shortcut Model policies in state and visual manipulation tasks achieved an average net increase of 40.34% after fine-tuning with ReinFlow at four or even one denoising step, whose performance is comparable to fine-tuned DDIM policies while saving computation time for an average of 23.20% . Code, model, and checkpoints available on the project website: https://reinflow.github.io/ * Corresponding author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
E π +∞ h=0 γ h r h (a h , o h ) . The Q function value function and the advantage function are defined as
(1) We drop the subscript h for V π h , Q π h , and A π h when the policy and POMDP are stationary.
Flow Matching Models Flow matching [32] transforms random variables from one distribution p 0 to another p 1 with flow mappings ψ : [0, 1] × X → X , where X t := ψ t (X 0 ), t ∈ [0, 1]. This process is associated with an ODE:
Practitioners also sample t from the beta distribution [7] or the logit normal distribution [16]. [18] proposes "Shortcut Models" to further improve the generation quality of Rectified Flow at very few denoising steps by enforcing the velocity generated by two steps to align with that generated by a single step. During inference, we numerically solve the transport equation by integrating the learned velocity field:
where K is the number of denoising steps, 0 = t 0 < t 1 < . . . < t K-1 < 1 = t K are the discretized time steps, ∆t i = t i+1 -t i is the step size.
Flow Matching Policy When we instantiate a flow-matching model in the action-generation setting, we obtain a flow-matching policy for robot learning. We denote by a t h the denoised action at time t
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf2f66de-0a27-4e1e-9bb1-f44f5ac3a80eCited by top-tier papers13
- SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential ModelingYixian Zhang, Shu'ang Yu, Tonghe Zhang, Mo Guang et al.ICLR 2026 · 33 citations
- RFS: Reinforcement learning with Residual flow steering for dexterous manipulationEntong Su, Tyler Westenbroek, Anusha Nagabandi, Abhishek GuptaICLR 2026 · 13 citations
- D²PPO: Diffusion Policy Policy Optimization with Dispersive LossGuowei Zou, Weibing Li, Hejun Wu, Yukun Qian et al.AAAI 2026 · 3 citations
- Generative Online Reinforcement LearningChubin Zhang, Zhenglin Wan, Feng Chen, Fuchao Yang et al.ICML 2026 · 1 citation
- Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot ManipulationQinglun Zhang, Shen Cheng, Tian Dan, Haoqiang Fan et al.CVPR 2026 · 1 citation
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 1,115 citations
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li et al.NeurIPS 2025 · 647 citations
- Data Quality in Imitation LearningSuneel Belkhale, Yuchen Cui, Dorsa SadighNeurIPS 2023 · 135 citations
Related papers
- Mean Flow Policy OptimizationXiaoyi Dong, Xi Zhang, Jian ChengICML 2026
- Flow Matching with Injected Noise for Offline-to-Online Reinforcement LearningYongjae Shin, Jongseong Chae, Jongeui Park, Youngchul SungICLR 2026 · 1 citation
- One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement LearningXuan Thanh Nguyen, Chang Dong YooICLR 2026 · 11 citations
- Flow Q-LearningSeohong Park, Qiyang Li, Sergey LevineICML 2025
- Flow Matching Policy GradientsDavid McAllister, Songwei Ge, Brent Yi, Chung Min Kim et al.ICLR 2026 · 103 citations
