Lune

NeurIPS2025Top-tier venue

DisCO: Reinforcing Large Reasoning Models with Discriminative Constrained Optimization

Gang Li, Ming Lin, Tomer Galanti, Zhengzhong Tu, Tianbao Yang

2025Year
24Citations
3Top-tier citations

Abstract

The recent success and openness of DeepSeek-R1 have brought widespread attention to Group Relative Policy Optimization (GRPO) as a reinforcement learning method for large reasoning models (LRMs). In this work, we analyze the GRPO objective under a binary reward setting and reveal an inherent limitation of questionlevel difficulty bias arising from its group relative advantage function. We also identify a connection between GRPO and traditional discriminative methods in supervised learning. Motivated by these insights, we introduce a new Discriminative Constrained Optimization (DisCO) framework for reinforcing LRMs, grounded in the principle of discriminative learning: increasing the scores of positive answers while decreasing those of negative ones. The main differences between DisCO and GRPO and its recent variants are: (1) it replaces the group relative objective with a discriminative objective defined by a scoring function; (2) it abandons clipping-based surrogates in favor of non-clipping RL surrogate objectives used as scoring functions; (3) it employs a simple yet effective constrained optimization approach to enforce the KL divergence constraint. As a result, DisCO offers notable advantages over GRPO and its variants: (i) it completely eliminates difficulty bias by adopting discriminative objectives; (ii) it addresses the entropy instability in GRPO and its variants through the use of non-clipping scoring functions and a constrained optimization approach, yielding long and stable training dynamics; (iii) it allows the incorporation of advanced discriminative learning techniques to address data imbalance, where a significant number of questions have more negative than positive generated answers during training. Our experiments on enhancing the mathematical reasoning capabilities of SFT-finetuned models show that DisCO significantly outperforms GRPO and its improved variants such as DAPO, achieving average gains of 7% over GRPO and 6% over DAPO across six benchmark tasks for a 1.5B model. 1 How can we design more effective optimization methods for reinforcing large reasoning models in a principled manner without inheriting the limitations of GRPO?

This paper addresses the above question through a complete redesign of the objective function, grounded in the principles of discriminative learning. Specifically, we first analyze the objective function of GRPO and its variants under a binary reward setting, leading to two key insights: (1) the root cause of GRPO's difficulty bias lies in its group relative advantage function, which induces disproportionately small weights to questions that are either too easy or too hard; and (2) there exists a conceptual connection to traditional discriminative approaches in AUC maximization, which aim to increase the scores of positive outputs while decreasing that of negative outputs.

Building upon these insights, we propose a principled optimization framework for reinforcing large reasoning models based on discriminative learning. Specifically, we optimize a discriminative objective using a proper scoring function over input-output pairs, which increases the score of positive outputs and decreases that of negative ones. The flexibility of our framework allows us to leverage simple non-clipping RL surrogate objectives as scoring functions without suffering from entropy instability, and to incorporate advanced discriminative techniques to address data imbalance in generated rollouts. To ensure training stability, we adopt a simple yet effective constrained optimization method to enforce a trust region constraint bounding the KL divergence between the updated model and the old model. Our experiments for mathematical reasoning show that DisCO significantly outperforms all baselines for fine-tuning DeepSeek-R1-Distill-Qwen and -Llama models with a maximum 8k response length for both training and inference, and also achieves a better performance than GRPO that uses a maximum 24k length for training and 32k length for inference.

Our main contributions are summarized as follows:

• We present an analysis of GRPO's objective function, identifying the root cause of difficulty bias and revealing its conceptual connection to classic discriminative methods for AUC maximization.

• We introduce a principled discriminative constrained optimization framework for reinforcing large reasoning models, which avoids both difficulty bias and training instability. This framework gives rise to a family of methods we refer to as DisCO.

• We demonstrate significant improvements of our DisCO method over GRPO and four other baselines, including DAPO, through experiments for fine-tuning LRMs on mathematical reasoning tasks, with evaluations across six benchmarks.

2 Related Work Large Reasoning Models (LRMs). Recent advances of LRMs, such as OpenAI o1 [51], DeepSeek-R1 [23] and Kimi K1.5 [63], have demonstrated strong reasoning capability in solving complex tasks. Departing from earlier approach

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ac0862f2-bfe8-4148-9ccd-b2c840ef26e9

Cited by top-tier papers3

Ask how each one uses it

Builds on24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines