When Do Curricula Work?
Xiaoxia Wu, Ethan Dyer, Behnam Neyshabur
Abstract
Inspired by human learning, researchers have proposed ordering examples during training based on their difficulty. Both curriculum learning, exposing a network to easier examples early in training, and anti-curriculum learning, showing the most difficult examples first, have been suggested as improvements to the standard i.i.d. training. In this work, we set out to investigate the relative benefits of ordered learning. We first investigate the implicit curricula resulting from architectural and optimization bias and find that samples are learned in a highly consistent order. Next, to quantify the benefit of explicit curricula, we conduct extensive experiments over thousands of orderings spanning three kinds of learning: curriculum, anti-curriculum, and random-curriculum -in which the size of the training dataset is dynamically increased over time, but the examples are randomly ordered. We find that for standard benchmark datasets, curricula have only marginal benefits, and that randomly ordered samples perform as well or better than curricula and anti-curricula, suggesting that any benefit is entirely due to the dynamic training set size. Inspired by common use cases of curriculum learning in practice, we investigate the role of limited training time budget and noisy data in the success of curriculum learning. Our experiments demonstrate that curriculum, but not anti-curriculum can indeed improve the performance either with limited training time budget or in existence of noisy data. * Work performed in part while Xiaoxia Wu was interning at Google. 1 Both GPT-3 and T5 have non-uniform mixing strategies. See (Brown et al., 2020, Section 2.2, Table 2.2) and (Raffel et al., 2019, Section 3.5) for more details.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 421579ce-f76d-40df-b5a8-d7239b26eaeaCited by top-tier papers39
- Deep Learning Through the Lens of Example DifficultyRobert J. N. Baldock, Hartmut Maennel, Behnam NeyshaburNeurIPS 2021 · 204 citations
- Skill-it! A data-driven skills framework for understanding and training language modelsMayee F. Chen, Nicholas Roberts, Kush Bhatia, Jue Wang et al.NeurIPS 2023 · 143 citations
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsAngelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt et al.ICLR 2024 · 119 citations
- ACPL: Anti-curriculum Pseudo-labelling for Semi-supervised Medical Image ClassificationFengbei Liu, Yu Tian, Yuanhong Chen, Yuyuan Liu et al.CVPR 2022 · 119 citations
- Self-Paced Contrastive Learning for Semi-supervised Medical Image Segmentation with Meta-labelsJizong Peng, Ping Wang, Christian Desrosiers, Marco PedersoliNeurIPS 2021 · 80 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Dynamic Curriculum Learning for Imbalanced Data ClassificationYiru Wang, Weihao Gan, Jie Yang, Wei Wu et al.ICCV 2019 · 263 citations
- Beyond Synthetic Noise: Deep Learning on Controlled Noisy LabelsLu Jiang, Di Huang, Mason Liu, Weilong YangICML 2020 · 241 citations
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task SemanticsVinay Venkatesh Ramasesh, Ethan Dyer, Maithra RaghuICLR 2021 · 207 citations
- Curriculum Loss: Robust Learning and Generalization against Label CorruptionYueming Lyu, Ivor W. TsangICLR 2020 · 190 citations
Related papers
- Curriculum Learning for Natural Language UnderstandingBenfeng Xu, Licheng Zhang, Zhendong Mao, Quan Wang et al.ACL 2020 · 156 citations
- Robust Curriculum Learning: from clean label detection to noisy label self-correctionTianyi Zhou, Shengjie Wang, Jeff A. BilmesICLR 2021 · 111 citations
- CurBench: Curriculum Learning BenchmarkYuwei Zhou, Zirui Pan, Xin Wang, Hong Chen et al.ICML 2024 · 11 citations
- Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLUFenia Christopoulou, Gerasimos Lampouras, Ignacio IacobacciEMNLP 2022 · 3 citations
- What Makes a Good Curriculum? Disentangling the Effects of Data Ordering on LLM Mathematical ReasoningYaning Jia, Chunhui Zhang, Xingjian Diao, Xiangchi Yuan et al.ACL 2026 · 4 citations
