Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
Deyang Kong, Qi Guo, Xiangyu Xi, Wei Wang, Jingang Wang, Xunliang Cai, Shikun Zhang, Wei Ye
摘要
The low sampling efficiency during the rollout phase poses a significant challenge to scaling reinforcement learning for large language model reasoning. Existing methods attempt to improve efficiency by scheduling problems based on problem difficulties. However, these approaches suffer from unstable and biased estimations of problem difficulty and fail to capture the alignment between model competence and problem difficulty in RL training, leading to suboptimal performance. To address these challenges, we introduce Competence-Difficulty Alignment Sampling (CDAS). This approach allows for accurate and stable estimation of problem difficulties by aggregating historical performance discrepancies across problems. Subsequently, model competence is quantified to adaptively select problems whose difficulties align with the model's current competence using a fixed-point system. Extensive experiments in mathematical RL training show that CDAS consistently outperforms strong baselines, achieving the highest average accuracy of 45.89%. Furthermore, CDAS reduces the training step time overhead by 57.06% compared to the widely-used Dynamic Sampling strategy, verifying the efficiency of CDAS. Additional experiments on different tasks, model architectures, and model sizes demonstrate the generalization capability of CDAS.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Adaptive Rollout Allocation for Online Reinforcement Learning with Verifiable RewardsHieu Trung Nguyen, Bao Nguyen, Wenao Ma, Yuzhi Zhao 等ICLR 2026 · 被引用 19 次
- Resource-Efficient Reinforcement for Reasoning Large Language Models via Dynamic One-Shot Policy RefinementYunjian Zhang, Sudong Wang, Yang Li, Peiran Xu 等ICML 2026 · 被引用 4 次
- Diffuse Thinking: Exploring Diffusion Language Models as Efficient Thought Proposers for ReasoningChenyang Shao, Sijian Ren, Fengli Xu, Yong LiACL 2026 · 被引用 4 次
- Counteracting the Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancingXin Guo, Zhiheng Xi, Yiwen Ding, Yitao Zhai 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper9
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- GPG: A Simple and Strong Reinforcement Learning Baseline for Model ReasoningXiangxiang Chu, Hailang Huang, Xiao Zhang, Fei Wei 等ICLR 2026 · 被引用 168 次
- HybridFlow: A Flexible and Efficient RLHF FrameworkGuangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu 等EuroSys 2025 · 被引用 61 次
相关 Paper
- HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget ReallocationFeng Xiong, Hongling Xu, Yifei Wang, Runxi Cheng 等EMNLP 2025 · 被引用 18 次
- Enhancing Efficiency and Exploration in Reinforcement Learning for LLMsMengqi Liao, Xiangyu Xi, Ruinian Chen, Jia Leng 等EMNLP 2025 · 被引用 16 次
- Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout ReplayYifan Sun, Jingyan Shen, Yibin Wang, Tianyu Chen 等NeurIPS 2025 · 被引用 63 次
- Tailoring the Training: Difficulty-Aware Learning Strategy Allocation for Large Language ModelsXiaoling Zhou, Shuaiyu Zhou, Zhemg Lee, Tao Chen 等ICML 2026
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning ModelsRunze Liu, Jiakang Wang, Yuling Shi, Zhihui Xie 等ICLR 2026 · 被引用 13 次
