What Makes a Good Curriculum? Disentangling the Effects of Data Ordering on LLM Mathematical Reasoning
Yaning Jia, Chunhui Zhang, Xingjian Diao, Xiangchi Yuan, Zhongyu Ouyang, Chiyu Ma, Soroush Vosoughi
Abstract
Curriculum learning (CL) - ordering training data from easy to hard - has become a popular strategy for improving reasoning in large language models (LLMs). Yet prior work employs disparate difficulty metrics and training setups, leaving open fundamental questions: When does curriculum help? Which direction - forward or reverse - is better? And does the answer depend on what we measure? We address these questions through a unified offline evaluation framework that decomposes curriculum difficulty into five complementary dimensions: Problem Difficulty, Model Surprisal, Confidence Margin, Predictive Uncertainty, and Decision Variability. Through controlled post-training experiments on mathematical reasoning benchmarks with Llama3.1-8B, Mistral-7B, and Gemma3-4B, we find that (i) no curriculum strategy dominates universally - the relative effectiveness of forward versus reverse CL depends jointly on model capability and task complexity; (ii) even within a single metric, samples at different difficulty levels produce distinct gains depending on task demands; and (iii) task-aligned curricula focus on shaping the model's final representations and generalization, whereas inner-state curricula modulate internal states such as confidence and uncertainty. Our findings challenge the notion of a universal curriculum strategy and offer actionable guidance across model and task regimes, with some metrics indicating that prioritizing decision-uncertain samples can further enhance learning outcomes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- D: Dynamic Directional Graph-Constrained Data Scheduling for LLM TrainingYuanjian Xu, Jianing Hao, Guang Zhang, Zhong LiICML 2026
- What Makes LLMs Effective Sequential Recommenders? A Study on Preference Intensity and Temporal ContextZhongyu Ouyang, Qianlong Wen, Chunhui Zhang, Yanfang Ye et al.ACL 2026
Builds on11
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu et al.ICLR 2024 · 637 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- Skill-it! A data-driven skills framework for understanding and training language modelsMayee F. Chen, Nicholas Roberts, Kush Bhatia, Jue Wang et al.NeurIPS 2023 · 143 citations
Related papers
- What Do Learning Dynamics Reveal About Generalization in LLM Mathematical Reasoning?Katie Kang, Amrith Setlur, Dibya Ghosh, Jacob Steinhardt et al.ICML 2025
- SLR: Automated Synthesis for Scalable Logical ReasoningLukas Helff, Ahmad Omar, Felix Friedrich, Antonia Wüst et al.ACL 2026 · 6 citations
- Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement LearningZhiheng Xi, Wenxiang Chen, Boyang Hong, Senjie Jin et al.ICML 2024 · 70 citations
- How do LLMs Compute Verbal Confidence?Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero et al.ICML 2026 · 16 citations
- On the Generalization Gap in Self-Evolving Language Model ReasoningZhenting Qi, Susanna Maria Baby, Stefanie Baby, Kan Yuan et al.ICML 2026
