Provable Advantage of Curriculum Learning on Parity Targets with Mixed Inputs
Emmanuel Abbe, Elisabetta Cornacchia, Aryo Lotfi
Abstract
Experimental results have shown that curriculum learning, i.e., presenting simpler examples before more complex ones, can improve the efficiency of learning. Some recent theoretical results also showed that changing the sampling distribution can help neural networks learn parities, with formal results only for large learning rates and one-step arguments. Here we show a separation result in the number of training steps with standard (bounded) learning rates on a common sample distribution: if the data distribution is a mixture of sparse and dense inputs, there exists a regime in which a 2-layer ReLU neural network trained by a curriculum noisy-GD (or SGD) algorithm that uses sparse examples first, can learn parities of sufficiently large degree, while any fully connected neural network of possibly larger width or depth trained by noisy-GD on the unordered samples cannot learn without additional steps. We also provide experimental results supporting the qualitative separation beyond the specific regime of the theoretical results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fdc71cb-f59f-4cae-8f0d-2b508dc3e25dCited by top-tier papers17
- Generalization on the Unseen, Logic Reasoning and Degree CurriculumEmmanuel Abbe, Samy Bengio, Aryo Lotfi, Kevin RizkICML 2023 · 68 citations
- How Far Can Transformers Reason? The Globality Barrier and Inductive ScratchpadEmmanuel Abbe, Samy Bengio, Aryo Lotfi, Colin Sandon et al.NeurIPS 2024 · 52 citations
- RL for Reasoning by Adaptively Revealing RationalesMohammad Hossein Amani, Aryo Lotfi, Nicolas Baldwin, Samy Bengio et al.ICLR 2026 · 19 citations
- Deep sequence models tend to memorize geometrically; it is unclear whyShahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv KumarICML 2026 · 11 citations
- Tilting the Odds at the Lottery: the Interplay of Overparameterisation and Curricula in Neural NetworksStefano Sarao Mannelli, Yaraslau Ivashinka, Andrew M. Saxe, Luca SagliettiICML 2024 · 11 citations
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade et al.NeurIPS 2022 · 220 citations
- Learning Parities with Neural NetworksAmit Daniely, Eran MalachNeurIPS 2020 · 104 citations
- Generalization on the Unseen, Logic Reasoning and Degree CurriculumEmmanuel Abbe, Samy Bengio, Aryo Lotfi, Kevin RizkICML 2023 · 68 citations
- Quantifying the Benefit of Using Differentiable Learning over Tangent KernelsEran Malach, Pritish Kamath, Emmanuel Abbe, Nathan SrebroICML 2021 · 44 citations
Related papers
- A Mathematical Model for Curriculum Learning for ParitiesElisabetta Cornacchia, Elchanan MosselICML 2023
- Learning High-Degree Parities: The Crucial Role of the InitializationEmmanuel Abbe, Elisabetta Cornacchia, Jan Hazla, Donald Kougang-YombiICLR 2025
- Feature learning via mean-field Langevin dynamics: classifying sparse parities and beyondTaiji Suzuki, Denny Wu, Kazusato Oko, Atsushi NitandaNeurIPS 2023 · 17 citations
- Learning High-Dimensional Parity Functions with Product Networks using Gradient DescentGuillaume Larue, Louis-Adrien Dufrène, Quentin Lampin, Hadi Ghauch et al.ICML 2026
- Understanding the effects of data parallelism and sparsity on neural network trainingNamhoon Lee, Thalaiyasingam Ajanthan, Philip H. S. Torr, Martin JaggiICLR 2021 · 8 citations
