When Does Curriculum Learning Help? A Theoretical Perspective
Raman Arora, Yunjuan Wang, Kaibo Zhang
Abstract
Curriculum learning has emerged as an effective strategy to enhance the training efficiency and generalization of machine learning models. However, its theoretical underpinnings remain relatively underexplored. In this work, we develop a theoretical framework for curriculum learning based on biased regularized empirical risk minimization (RERM), identifying conditions under which curriculum learning provably improves generalization. We introduce a sufficient condition that characterizes a "good" curriculum and analyze a multi-task curriculum framework, where solving a sequence of convex tasks can facilitate better generalization. We also demonstrate how these theoretical insights translate to practical benefits when using stochastic gradient descent (SGD) as an optimization method. Beyond convex settings, we explore the utility of curriculum learning for non-convex tasks. Empirical evaluations on synthetic datasets and MNIST validate our theoretical findings and highlight the practical efficacy of curriculum-based training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a109d5eb-b744-4c2a-b412-eac7e8f48912Builds on4
- Curriculum By SmoothingSamarth Sinha, Animesh Garg, Hugo LarochelleNeurIPS 2020 · 95 citations
- An Analytical Theory of Curriculum Learning in Teacher-Student NetworksLuca Saglietti, Stefano Sarao Mannelli, Andrew M. SaxeNeurIPS 2022 · 43 citations
- Provable Advantage of Curriculum Learning on Parity Targets with Mixed InputsEmmanuel Abbe, Elisabetta Cornacchia, Aryo LotfiNeurIPS 2023 · 29 citations
- On the Statistical Benefits of Curriculum LearningZiping Xu, Ambuj TewariICML 2022 · 12 citations
Related papers
- SGD: The Role of Implicit Regularization, Batch-size and Multiple-epochsAyush Sekhari, Karthik Sridharan, Satyen KaleNeurIPS 2021 · 36 citations
- Understanding the Complexity Gains of Single-Task RL with a CurriculumQiyang Li, Yuexiang Zhai, Yi Ma, Sergey LevineICML 2023 · 21 citations
- Adaptive Curriculum LearningYajing Kong, Liu Liu, Jun Wang, Dacheng TaoICCV 2021 · 61 citations
- Investigating the Role of Weight Decay in Enhancing Nonconvex SGDTao Sun, Yuhao Huang, Li Shen, Kele Xu et al.CVPR 2025
- CurBench: Curriculum Learning BenchmarkYuwei Zhou, Zirui Pan, Xin Wang, Hong Chen et al.ICML 2024 · 11 citations
