ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
Tonghe Zhang, Chao Yu, Sichang Su, Yu Wang
摘要
We propose ReinFlow, a simple yet effective online reinforcement learning (RL) framework that fine-tunes a family of flow matching policies for continuous robotic control. Derived from rigorous RL theory, ReinFlow injects learnable noise into a flow policy's deterministic path, converting the flow into a discrete-time Markov Process for exact and straightforward likelihood computation. This conversion facilitates exploration and ensures training stability, enabling ReinFlow to fine-tune diverse flow model variants stably, including Rectified Flow [34] and Shortcut Models [18], particularly at very few or even one denoising step. We benchmark ReinFlow in representative locomotion and manipulation tasks, including longhorizon planning with visual input and sparse reward. The episode reward of Rectified Flow policies obtained an average net growth of 135.36% after fine-tuning in challenging legged locomotion tasks while saving denoising steps and 82.63% of wall time compared to state-of-the-art diffusion RL fine-tuning method DPPO [42]. The success rate of the Shortcut Model policies in state and visual manipulation tasks achieved an average net increase of 40.34% after fine-tuning with ReinFlow at four or even one denoising step, whose performance is comparable to fine-tuned DDIM policies while saving computation time for an average of 23.20% . Code, model, and checkpoints available on the project website: https://reinflow.github.io/ * Corresponding author 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
E π +∞ h=0 γ h r h (a h , o h ) . The Q function value function and the advantage function are defined as
(1) We drop the subscript h for V π h , Q π h , and A π h when the policy and POMDP are stationary.
Flow Matching Models Flow matching [32] transforms random variables from one distribution p 0 to another p 1 with flow mappings ψ : [0, 1] × X → X , where X t := ψ t (X 0 ), t ∈ [0, 1]. This process is associated with an ODE:
Practitioners also sample t from the beta distribution [7] or the logit normal distribution [16]. [18] proposes "Shortcut Models" to further improve the generation quality of Rectified Flow at very few denoising steps by enforcing the velocity generated by two steps to align with that generated by a single step. During inference, we numerically solve the transport equation by integrating the learned velocity field:
where K is the number of denoising steps, 0 = t 0 < t 1 < . . . < t K-1 < 1 = t K are the discretized time steps, ∆t i = t i+1 -t i is the step size.
Flow Matching Policy When we instantiate a flow-matching model in the action-generation setting, we obtain a flow-matching policy for robot learning. We denote by a t h the denoised action at time t
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential ModelingYixian Zhang, Shu'ang Yu, Tonghe Zhang, Mo Guang 等ICLR 2026 · 被引用 33 次
- RFS: Reinforcement learning with Residual flow steering for dexterous manipulationEntong Su, Tyler Westenbroek, Anusha Nagabandi, Abhishek GuptaICLR 2026 · 被引用 13 次
- D²PPO: Diffusion Policy Policy Optimization with Dispersive LossGuowei Zou, Weibing Li, Hejun Wu, Yukun Qian 等AAAI 2026 · 被引用 3 次
- Generative Online Reinforcement LearningChubin Zhang, Zhenglin Wan, Feng Chen, Fuchao Yang 等ICML 2026 · 被引用 1 次
- Efficient Hybrid SE(3)-Equivariant Visuomotor Flow Policy via Spherical Harmonics for Robot ManipulationQinglun Zhang, Shen Cheng, Tian Dan, Haoqiang Fan 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 被引用 1,115 次
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li 等NeurIPS 2025 · 被引用 647 次
- Data Quality in Imitation LearningSuneel Belkhale, Yuchen Cui, Dorsa SadighNeurIPS 2023 · 被引用 135 次
相关 Paper
- Mean Flow Policy OptimizationXiaoyi Dong, Xi Zhang, Jian ChengICML 2026
- Flow Matching with Injected Noise for Offline-to-Online Reinforcement LearningYongjae Shin, Jongseong Chae, Jongeui Park, Youngchul SungICLR 2026 · 被引用 1 次
- One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement LearningXuan Thanh Nguyen, Chang Dong YooICLR 2026 · 被引用 11 次
- Flow Q-LearningSeohong Park, Qiyang Li, Sergey LevineICML 2025
- Flow Matching Policy GradientsDavid McAllister, Songwei Ge, Brent Yi, Chung Min Kim 等ICLR 2026 · 被引用 103 次
