Lune

NeurIPS2025Top-tier venue

Q#: Provably Optimal Distributional RL for LLM Post-Training

Jin Peng Zhou, Kaiwen Wang, Jonathan D. Chang, Zhaolin Gao, Nathan Kallus, Kilian Q. Weinberger, Kianté Brantley, Wen Sun

2025Year
18Citations
1Top-tier citations

Abstract

Reinforcement learning (RL) post-training is crucial for LLM alignment and reasoning, but existing policy-based methods, such as PPO and DPO, can fall short of fixing shortcuts inherited from pre-training. In this work, we introduce Q♯Q\sharp, a value-based algorithm for KL-regularized RL that guides the reference policy using the optimal regularized QQ function. We propose to learn the optimal QQ function using distributional RL on an aggregated online dataset. Unlike prior value-based baselines that guide the model using unregularized QQ-values, our method is theoretically principled and provably learns the optimal policy for the KL-regularized RL problem. Empirically, Q♯Q\sharp outperforms prior baselines in math reasoning benchmarks while maintaining a smaller KL divergence to the reference policy. Theoretically, we establish a reduction from KL-regularized RL to no-regret online learning, providing the first bounds for deterministic MDPs under only realizability. Thanks to distributional RL, our bounds are also variance-dependent and converge faster when the reference policy has small variance. In sum, our results highlight Q♯Q\sharp as an effective approach for post-training LLMs, offering both improved performance and theoretical guarantees. The code can be found at https://github.com/jinpz/q_sharp.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 7b632871-7b8b-440f-83a0-d01ba610619e

Cited by top-tier papers1

Ask how each one uses it

Builds on35

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines