Flow-Based Policy for Online Reinforcement Learning
Lei Lyu, Yunfei Li, Yu Luo, Fuchun Sun, Tao Kong, Jiafeng Xu, Xiao Ma
Abstract
We present FlowRL, a novel framework for online reinforcement learning that integrates flow-based policy representation with Wasserstein-2-regularized optimization. We argue that in addition to training signals, enhancing the expressiveness of the policy class is crucial for the performance gains in RL. Flow-based generative models offer such potential, excelling at capturing complex, multimodal action distributions. However, their direct application in online RL is challenging due to a fundamental objective mismatch: standard flow training optimizes for static data imitation, while RL requires value-based policy optimization through a dynamic buffer, leading to difficult optimization landscapes. FlowRL first models policies via a state-dependent velocity field, generating actions through deterministic ODE integration from noise. We derive a constrained policy search objective that jointly maximizes Q through the flow policy while bounding the Wasserstein-2 distance to a behavior-optimal policy implicitly derived from the replay buffer. This formulation effectively aligns the flow optimization with the RL objective, enabling efficient and value-aware policy learning despite the complexity of the policy class. Empirical evaluations on DMControl and Humanoidbench demonstrate that FlowRL achieves competitive performance in online reinforcement learning benchmarks.We have released our code here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86eb7937-9302-46b6-92d7-7a0c8b98aaf1Cited by top-tier papers13
- FlowRL: Matching Reward Distributions for LLM ReasoningXuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li et al.ICLR 2026 · 41 citations
- SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential ModelingYixian Zhang, Shu'ang Yu, Tonghe Zhang, Mo Guang et al.ICLR 2026 · 33 citations
- Scalable Exploration for High-Dimensional Continuous Control via Value-Guided FlowYunyue Wei, Chenhui Zuo, Yanan SuiICLR 2026 · 8 citations
- Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow PoliciesZeyang Li, Sunbochen Tang, Navid AzizanICML 2026 · 6 citations
- FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge MatchingLei Lyu, Yunfei Li, Yu Luo, Fuchun Sun et al.ICML 2026 · 5 citations
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- Mean Flow Policy OptimizationXiaoyi Dong, Xi Zhang, Jian ChengICML 2026
- Generative Online Reinforcement LearningChubin Zhang, Zhenglin Wan, Feng Chen, Fuchao Yang et al.ICML 2026 · 1 citation
- Flow Q-LearningSeohong Park, Qiyang Li, Sergey LevineICML 2025
- Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein RegularizationJiajun Fan, Shuaike Shen, Chaoran Cheng, Yuxin Chen et al.ICLR 2025
- Behavior Regularization with Flow Latent Policy for Offline Reinforcement LearningYulong Xia, Fuchun SunAAAI 2026
