When Do Curricula Work?
Xiaoxia Wu, Ethan Dyer, Behnam Neyshabur
摘要
Inspired by human learning, researchers have proposed ordering examples during training based on their difficulty. Both curriculum learning, exposing a network to easier examples early in training, and anti-curriculum learning, showing the most difficult examples first, have been suggested as improvements to the standard i.i.d. training. In this work, we set out to investigate the relative benefits of ordered learning. We first investigate the implicit curricula resulting from architectural and optimization bias and find that samples are learned in a highly consistent order. Next, to quantify the benefit of explicit curricula, we conduct extensive experiments over thousands of orderings spanning three kinds of learning: curriculum, anti-curriculum, and random-curriculum -in which the size of the training dataset is dynamically increased over time, but the examples are randomly ordered. We find that for standard benchmark datasets, curricula have only marginal benefits, and that randomly ordered samples perform as well or better than curricula and anti-curricula, suggesting that any benefit is entirely due to the dynamic training set size. Inspired by common use cases of curriculum learning in practice, we investigate the role of limited training time budget and noisy data in the success of curriculum learning. Our experiments demonstrate that curriculum, but not anti-curriculum can indeed improve the performance either with limited training time budget or in existence of noisy data. * Work performed in part while Xiaoxia Wu was interning at Google. 1 Both GPT-3 and T5 have non-uniform mixing strategies. See (Brown et al., 2020, Section 2.2, Table 2.2) and (Raffel et al., 2019, Section 3.5) for more details.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Deep Learning Through the Lens of Example DifficultyRobert J. N. Baldock, Hartmut Maennel, Behnam NeyshaburNeurIPS 2021 · 被引用 204 次
- Skill-it! A data-driven skills framework for understanding and training language modelsMayee F. Chen, Nicholas Roberts, Kush Bhatia, Jue Wang 等NeurIPS 2023 · 被引用 143 次
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsAngelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt 等ICLR 2024 · 被引用 119 次
- ACPL: Anti-curriculum Pseudo-labelling for Semi-supervised Medical Image ClassificationFengbei Liu, Yu Tian, Yuanhong Chen, Yuyuan Liu 等CVPR 2022 · 被引用 119 次
- Self-Paced Contrastive Learning for Semi-supervised Medical Image Segmentation with Meta-labelsJizong Peng, Ping Wang, Christian Desrosiers, Marco PedersoliNeurIPS 2021 · 被引用 80 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Dynamic Curriculum Learning for Imbalanced Data ClassificationYiru Wang, Weihao Gan, Jie Yang, Wei Wu 等ICCV 2019 · 被引用 263 次
- Beyond Synthetic Noise: Deep Learning on Controlled Noisy LabelsLu Jiang, Di Huang, Mason Liu, Weilong YangICML 2020 · 被引用 241 次
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task SemanticsVinay Venkatesh Ramasesh, Ethan Dyer, Maithra RaghuICLR 2021 · 被引用 207 次
- Curriculum Loss: Robust Learning and Generalization against Label CorruptionYueming Lyu, Ivor W. TsangICLR 2020 · 被引用 190 次
相关 Paper
- Curriculum Learning for Natural Language UnderstandingBenfeng Xu, Licheng Zhang, Zhendong Mao, Quan Wang 等ACL 2020 · 被引用 156 次
- Robust Curriculum Learning: from clean label detection to noisy label self-correctionTianyi Zhou, Shengjie Wang, Jeff A. BilmesICLR 2021 · 被引用 111 次
- CurBench: Curriculum Learning BenchmarkYuwei Zhou, Zirui Pan, Xin Wang, Hong Chen 等ICML 2024 · 被引用 11 次
- Training Dynamics for Curriculum Learning: A Study on Monolingual and Cross-lingual NLUFenia Christopoulou, Gerasimos Lampouras, Ignacio IacobacciEMNLP 2022 · 被引用 3 次
- What Makes a Good Curriculum? Disentangling the Effects of Data Ordering on LLM Mathematical ReasoningYaning Jia, Chunhui Zhang, Xingjian Diao, Xiangchi Yuan 等ACL 2026 · 被引用 4 次
