FlowRL: Matching Reward Distributions for LLM Reasoning
Xuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li, Kaiyan Zhang, Che Jiang, Youbang Sun, Ermo Hua, Yuxin Zuo, Xingtai Lv, Qizheng Zhang, Lin Chen
Abstract
We propose FlowRL: matching the full reward distribution via flow balancing instead of solely maximizing rewards in large language model (LLM) reinforcement learning (RL). Recent advanced reasoning models adopt reward-maximizing methods (e.g., PPO and GRPO), which tend to over-optimize dominant reward signals while neglecting less frequent but valid reasoning paths, thus reducing diversity. In contrast, we transform scalar rewards into a normalized target distribution using a learnable partition function, and then minimize the reverse KL divergence between the policy and the target distribution. We implement this idea as a flow-balanced optimization method that promotes diverse exploration and generalizable reasoning trajectories. We conduct experiments on both math and code reasoning tasks: FlowRL achieves a significant average improvement of over GRPO and over PPO on math benchmarks, and performs consistently better on code reasoning tasks. These results highlight reward distribution-matching as a key step toward efficient exploration and diverse reasoning in LLM reinforcement learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a42fceb8-d4e5-415d-af60-83a6d5f381abCited by top-tier papers10
- Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)Zhenjie Yang, Xiaosong Jia, Qifeng Li, Xue Yang et al.NeurIPS 2025 · 65 citations
- Enhancing Reasoning for Diffusion LLMs via Distribution Matching Policy OptimizationYuchen Zhu, Wei Guo, Jaemoo Choi, Petr Molodyk et al.ICML 2026 · 13 citations
- Efficient Paths and Dense Rewards: Probabilistic Flow Reasoning for Large Language ModelsYan Liu, Feng Zhang, Zhanyu Ma, Jun Xu et al.ACL 2026 · 2 citations
- Beyond Mode Collapse: Distribution Matching for Diverse ReasoningXiaozhe Li, Yang Li, Xinyu Fang, Shengyuan Ding et al.ICML 2026 · 2 citations
- Expected Return Causes Outcome-Level Mode Collapse in Reinforcement Learning and How to Fix It with Inverse Probability ScalingAbhijeet Sinha, Sundari Elango, Dianbo LiuICML 2026 · 2 citations
Builds on39
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 1,126 citations
Related papers
- R1-ShareVL: Incentivizing Reasoning Capabilities of Multimodal Large Language Models via Share-GRPOHuanjin Yao, Qixiang Yin, Jingyi Zhang, Min Yang et al.NeurIPS 2025 · 3 citations
- Random Policy Valuation is Enough for LLM Reasoning with Verifiable RewardsHaoran He, Yuxiao Ye, Qingpeng Cai, Chen Hu et al.ICLR 2026 · 9 citations
- Risk-Sensitive Reinforcement Learning for Alleviating Exploration Dilemmas in Large Language ModelsYuhua Jiang, Jiawei Huang, Yufeng Yuan, Xin Mao et al.ICLR 2026 · 8 citations
- RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-TrainingTao Ren, Jinyang Jiang, Hui Yang, Wan Tian et al.ICLR 2026 · 8 citations
- XRPO: Pushing the Limits of GRPO with Targeted Exploration and ExploitationUdbhav Bamba, Minghao Fang, Yifan Yu, Haizhong Zheng et al.ICML 2026 · 17 citations
