Lune

NeurIPS2025顶会

DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization

Gang Li, Ming Lin, Tomer Galanti, Zhengzhong Tu, Tianbao Yang

2025年份
24被引次数
3顶会引用

摘要

The recent success and openness of DeepSeek-R1 have brought widespread attention to Group Relative Policy Optimization (GRPO) as a reinforcement learning method for large reasoning models (LRMs). In this work, we analyze the GRPO objective under a binary reward setting and reveal an inherent limitation of questionlevel difficulty bias arising from its group relative advantage function. We also identify a connection between GRPO and traditional discriminative methods in supervised learning. Motivated by these insights, we introduce a new Discriminative Constrained Optimization (DisCO) framework for reinforcing LRMs, grounded in the principle of discriminative learning: increasing the scores of positive answers while decreasing those of negative ones. The main differences between DisCO and GRPO and its recent variants are: (1) it replaces the group relative objective with a discriminative objective defined by a scoring function; (2) it abandons clipping-based surrogates in favor of non-clipping RL surrogate objectives used as scoring functions; (3) it employs a simple yet effective constrained optimization approach to enforce the KL divergence constraint. As a result, DisCO offers notable advantages over GRPO and its variants: (i) it completely eliminates difficulty bias by adopting discriminative objectives; (ii) it addresses the entropy instability in GRPO and its variants through the use of non-clipping scoring functions and a constrained optimization approach, yielding long and stable training dynamics; (iii) it allows the incorporation of advanced discriminative learning techniques to address data imbalance, where a significant number of questions have more negative than positive generated answers during training. Our experiments on enhancing the mathematical reasoning capabilities of SFT-finetuned models show that DisCO significantly outperforms GRPO and its improved variants such as DAPO, achieving average gains of 7% over GRPO and 6% over DAPO across six benchmark tasks for a 1.5B model. 1 How can we design more effective optimization methods for reinforcing large reasoning models in a principled manner without inheriting the limitations of GRPO?

This paper addresses the above question through a complete redesign of the objective function, grounded in the principles of discriminative learning. Specifically, we first analyze the objective function of GRPO and its variants under a binary reward setting, leading to two key insights: (1) the root cause of GRPO's difficulty bias lies in its group relative advantage function, which induces disproportionately small weights to questions that are either too easy or too hard; and (2) there exists a conceptual connection to traditional discriminative approaches in AUC maximization, which aim to increase the scores of positive outputs while decreasing that of negative outputs.

Building upon these insights, we propose a principled optimization framework for reinforcing large reasoning models based on discriminative learning. Specifically, we optimize a discriminative objective using a proper scoring function over input-output pairs, which increases the score of positive outputs and decreases that of negative ones. The flexibility of our framework allows us to leverage simple non-clipping RL surrogate objectives as scoring functions without suffering from entropy instability, and to incorporate advanced discriminative techniques to address data imbalance in generated rollouts. To ensure training stability, we adopt a simple yet effective constrained optimization method to enforce a trust region constraint bounding the KL divergence between the updated model and the old model. Our experiments for mathematical reasoning show that DisCO significantly outperforms all baselines for fine-tuning DeepSeek-R1-Distill-Qwen and -Llama models with a maximum 8k response length for both training and inference, and also achieves a better performance than GRPO that uses a maximum 24k length for training and 32k length for inference.

Our main contributions are summarized as follows:

• We present an analysis of GRPO's objective function, identifying the root cause of difficulty bias and revealing its conceptual connection to classic discriminative methods for AUC maximization.

• We introduce a principled discriminative constrained optimization framework for reinforcing large reasoning models, which avoids both difficulty bias and training instability. This framework gives rise to a family of methods we refer to as DisCO.

• We demonstrate significant improvements of our DisCO method over GRPO and four other baselines, including DAPO, through experiments for fine-tuning LRMs on mathematical reasoning tasks, with evaluations across six benchmarks.

2 Related Work Large Reasoning Models (LRMs). Recent advances of LRMs, such as OpenAI o1 [51], DeepSeek-R1 [23] and Kimi K1.5 [63], have demonstrated strong reasoning capability in solving complex tasks. Departing from earlier approach

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

它引用的顶会 Paper24

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖