The Easy, the Hard, and the Learnable: Confidence and Difficulty-Adaptive Policy Optimization for LLM Reasoning
(Andrew) Zhanke Zhou, Xiangyu Lu, Chentao Cao, Brando Miranda, Tongliang Liu, Bo Han, Sanmi Koyejo
Abstract
RL with verifiable rewards can substantially improve LLM reasoning, yet standard GRPO-style training often treats easy, hard, and learnable questions alike through uniform sampling and weighting, leading to inefficient compute allocation. We study GRPO by tracking token logprobabilities, group-normalized advantages, and the induced token-level update weights. This reveals three recurring dynamics as training proceeds: (1) confidence inflation, (2) advantage contraction, and (3) hierarchical convergence. These findings suggest that the utility of each update depends strongly on both question difficulty and the model's current competence. Motivated by this, we propose Confidence and Difficulty-adaptive Policy Optimization (CoDaPO), which assigns each question a bounded value from rollout confidence and empirical difficulty. CoDaPO then uses this value to reweight policy updates and resample high-value learnable questions within mini-batches, thereby increasing discovery within the learnable band under a fixed compute budget. Across twelve benchmarks, CoDaPO consistently improves accuracy over existing RL methods. Our code is publicly available at https: //github.com/tmlr-group/CoDaPO .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45acfaef-eb5b-4fa5-9e64-978a7da911e5Builds on24
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
Related papers
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy EntropyHongze Tan, Zihan Wang, Jianfei Pan, Jinghao Lin et al.ICML 2026 · 53 citations
- RiskPO: Risk-based Policy Optimization with Verifiable Reward for LLM Post-TrainingTao Ren, Jinyang Jiang, Hui Yang, Wan Tian et al.ICLR 2026 · 8 citations
- Advantage Collapse in Group Relative Policy Optimization: Diagnosis and MitigationXixiang He, Qiyao Sun, Ao Cheng, Xingming Li et al.ICML 2026
- Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMsZhihe Yang, Xufang Luo, Zilong Wang, Dongqi Han et al.ICLR 2026 · 47 citations
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement LearningWenlong Deng, Yi Ren, Yushu Li, Boying Gong et al.ICLR 2026 · 9 citations
