Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation
Enci Zhang, Xingang Yan, Wei Lin, Tianxiang Zhang, Qianchun Lu
Abstract
Despite impressive progress in areas like mathematical reasoning, large language models still face significant challenges in consistently solving complex problems. Drawing inspiration from key human learning strategies, we propose two novel strategies to enhance the capability of large language models to solve these complex problems. First, Adaptive Difficulty Curriculum Learning (ADCL) is a novel curriculum learning strategy that tackles the Difficulty Shift phenomenon (i.e., a model's perception of problem difficulty dynamically changes during training) by periodically re-estimating difficulty within upcoming data batches to maintain alignment with the model's evolving capabilities. Second, Expert-Guided Self-Reformulation (EGSR) is a novel reinforcement learning strategy that bridges the gap between imitation learning and pure exploration by guiding models to reformulate expert solutions within their own conceptual framework, rather than relying on direct imitation, fostering deeper understanding and knowledge assimilation. Extensive experiments on challenging mathematical reasoning benchmarks, using Qwen2.5-7B as the base model, demonstrate that these human-inspired strategies synergistically and significantly enhance performance. Notably, their combined application improves performance over the standard Zero-RL baseline by 10% on the AIME24 benchmark and 16.6% on AIME25.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Multi-Task GRPO: Reliable LLM Reasoning Across TasksShyam Sundhar Ramesh, Xiaotong Ji, Matthieu Zimmer, Sangwoong Yoon et al.ICML 2026 · 8 citations
- Beyond Oracle: Verifier-Supervision for Instruction Hierarchy in Reasoning and Instruction-Tuned LLMsSian-Yao Huang, Li-Hsien Chang, Che-Yu Lin, Cheng-Lin YangNeurIPS 2025 · 4 citations
- Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical StudyZhiheng Xi, Xin Guo, Jiaqi Liu, Jiazheng Zhang et al.ICML 2026 · 3 citations
- From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language ModelsJuncheng Wu, Hardy Chen, Haoqin Tu, Xianfeng Tang et al.ICML 2026
- Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVRDohyung Kim, Minbeom Kim, Jeonghye Kim, Lee Sangmook et al.ICML 2026
Builds on4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
Related papers
- AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical RevisitingRenda Li, Hailang Huang, Fei Wei, Feng Xiong et al.AAAI 2026 · 1 citation
- Tailoring the Training: Difficulty-Aware Learning Strategy Allocation for Large Language ModelsXiaoling Zhou, Shuaiyu Zhou, Zhemg Lee, Tao Chen et al.ICML 2026
- UR² : Unify RAG and Reasoning through Reinforcement LearningWeitao Li, Boran Xiang, Xiaolong Wang, Jingyi Ren et al.ACL 2026 · 1 citation
- Adaption-of-Thought: Learning Question Difficulty Improves Large Language Models for ReasoningMayi Xu, Yongqi Li, Ke Sun, Tieyun QianEMNLP 2024 · 1 citation
- Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM ReasoningShubham Parashar, Shurui Gui, Xiner Li, Hongyi Ling et al.ICLR 2026 · 112 citations
