Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation
Enci Zhang, Xingang Yan, Wei Lin, Tianxiang Zhang, Qianchun Lu
摘要
Despite impressive progress in areas like mathematical reasoning, large language models still face significant challenges in consistently solving complex problems. Drawing inspiration from key human learning strategies, we propose two novel strategies to enhance the capability of large language models to solve these complex problems. First, Adaptive Difficulty Curriculum Learning (ADCL) is a novel curriculum learning strategy that tackles the Difficulty Shift phenomenon (i.e., a model's perception of problem difficulty dynamically changes during training) by periodically re-estimating difficulty within upcoming data batches to maintain alignment with the model's evolving capabilities. Second, Expert-Guided Self-Reformulation (EGSR) is a novel reinforcement learning strategy that bridges the gap between imitation learning and pure exploration by guiding models to reformulate expert solutions within their own conceptual framework, rather than relying on direct imitation, fostering deeper understanding and knowledge assimilation. Extensive experiments on challenging mathematical reasoning benchmarks, using Qwen2.5-7B as the base model, demonstrate that these human-inspired strategies synergistically and significantly enhance performance. Notably, their combined application improves performance over the standard Zero-RL baseline by 10% on the AIME24 benchmark and 16.6% on AIME25.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Multi-Task GRPO: Reliable LLM Reasoning Across TasksShyam Sundhar Ramesh, Xiaotong Ji, Matthieu Zimmer, Sangwoong Yoon 等ICML 2026 · 被引用 8 次
- Beyond Oracle: Verifier-Supervision for Instruction Hierarchy in Reasoning and Instruction-Tuned LLMsSian-Yao Huang, Li-Hsien Chang, Che-Yu Lin, Cheng-Lin YangNeurIPS 2025 · 被引用 4 次
- Does Reinforcement Fine-Tuning Improve Generalization of LLM Agents? An Empirical StudyZhiheng Xi, Xin Guo, Jiaqi Liu, Jiazheng Zhang 等ICML 2026 · 被引用 3 次
- From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language ModelsJuncheng Wu, Hardy Chen, Haoqin Tu, Xianfeng Tang 等ICML 2026
- Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVRDohyung Kim, Minbeom Kim, Jeonghye Kim, Lee Sangmook 等ICML 2026
它引用的顶会 Paper4
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
相关 Paper
- AdaCuRL: Adaptive Curriculum Reinforcement Learning with Invalid Sample Mitigation and Historical RevisitingRenda Li, Hailang Huang, Fei Wei, Feng Xiong 等AAAI 2026 · 被引用 1 次
- Tailoring the Training: Difficulty-Aware Learning Strategy Allocation for Large Language ModelsXiaoling Zhou, Shuaiyu Zhou, Zhemg Lee, Tao Chen 等ICML 2026
- UR² : Unify RAG and Reasoning through Reinforcement LearningWeitao Li, Boran Xiang, Xiaolong Wang, Jingyi Ren 等ACL 2026 · 被引用 1 次
- Adaption-of-Thought: Learning Question Difficulty Improves Large Language Models for ReasoningMayi Xu, Yongqi Li, Ke Sun, Tieyun QianEMNLP 2024 · 被引用 1 次
- Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM ReasoningShubham Parashar, Shurui Gui, Xiner Li, Hongyi Ling 等ICLR 2026 · 被引用 112 次
