DisPPO: Quantile-Based Distributional Reinforcement Learning for Large Language Models
Zhijian Zhou, Long Li, Xuan Zhang, Zongkai Liu, Yanting Miao, Yuchen Liu, Deshu Chen, Ke Li, Xing Sun, Ruoxi Jiang, Xiaoyu Tan, Chao Qu, Yuan Qi
Abstract
Reinforcement Learning (RL) has become a cornerstone for enhancing the reasoning capabilities of Large Language Models (LLMs). However, standard actor-critic methods, such as PPO, rely on scalar value functions that estimate only the expectation of cumulative returns. This reduction inherently discards higher-order statistical information (e.g., variance and multimodality), leading to inaccurate value estimation and suboptimal credit assignment in complex tasks. While Distributional RL offers a solution by modeling the full return distribution, its application to LLMs remains challenging due to the computational intractability of value-based operations over large vocabularies and the instability and memory burden of off-policy replay mechanisms. In this paper, we propose DisPPO, a novel on-policy framework that seamlessly integrates non-parametric quantile regression into PPO. Theoretically, we prove that our distributional update operator---composed of the -return Bellman operator and quantile projection---is a contraction mapping in the Wasserstein metric, guaranteeing convergence to a unique fixed point. Empirically, we evaluate DisPPO using Llama and Qwen models across diverse benchmarks, including mathematical reasoning and Text-to-SQL generation. DisPPO consistently outperforms standard PPO and recent group-based baselines in both Pass@1 and Pass@ metrics, demonstrating that distributional critics provide a richer, more robust learning signal for large-scale reasoning models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext caa998c9-6fdb-48a9-bbb4-2b9fff299da2Builds on7
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMsJiazhan Feng, Shijue Huang, Xingwei Qu, Ge Zhang et al.ICLR 2026 · 406 citations
- Reasoning with Exploration: An Entropy PerspectiveDaixuan Cheng, Shaohan Huang, Xuekai Zhu, Bo Dai et al.AAAI 2026 · 216 citations
- The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable RewardLong Li, Zhijian Zhou, Jiaran Hao, Jason Klein Liu et al.ICLR 2026 · 46 citations
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific ProblemsChaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu et al.ACL 2024 · 18 citations
Related papers
- Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoningJiashun Liu, Johan S. Obando-Ceron, Han Lu, Yancheng He et al.ICLR 2026 · 11 citations
- Q#: Provably Optimal Distributional RL for LLM Post-TrainingJin Peng Zhou, Kaiwen Wang, Jonathan D. Chang, Zhaolin Gao et al.NeurIPS 2025 · 18 citations
- DAPO : Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage-Based Policy OptimizationJiacai Liu, Chaojie Wang, Chris Yuhao Liu, Liang Zeng et al.NeurIPS 2025 · 9 citations
- FlowRL: Matching Reward Distributions for LLM ReasoningXuekai Zhu, Daixuan Cheng, Dinghuai Zhang, Hengli Li et al.ICLR 2026 · 41 citations
- Discounted Beta–Bernoulli Reward Estimation for Sample-Efficient Reinforcement Learning with Verifiable RewardsHaechan Kim, Soohyun Ryu, Gyouk Chu, Doohyuk Jang et al.ICML 2026
