Pareto Frontiers in Deep Feature Learning: Data, Compute, Width, and Luck
Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Eran Malach, Cyril Zhang
Abstract
In modern deep learning, algorithmic choices (such as width, depth, and learning rate) are known to modulate nuanced resource tradeoffs. This work investigates how these complexities necessarily arise for feature learning in the presence of computational-statistical gaps. We begin by considering offline sparse parity learning, a supervised classification problem which admits a statistical query lower bound for gradient-based training of a multilayer perceptron. This lower bound can be interpreted as a multi-resource tradeoff frontier : successful learning can only occur if one is sufficiently rich (large model), knowledgeable (large dataset), patient (many training iterations), or lucky (many random guesses). We show, theoretically and experimentally, that sparse initialization and increasing network width yield significant improvements in sample efficiency in this setting. Here, width plays the role of parallel search: it amplifies the probability of finding “lottery ticket” neurons, which learn sparse features more sample-efficiently. Finally, we show that the synthetic sparse parity task can be useful as a proxy for real problems requiring axis-aligned feature learning. We demonstrate improved sample efficiency on tabular classification benchmarks by using wide, sparsely-initialized MLP models; these networks sometimes outperform tuned random forests.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bad46bab-3bc2-4d14-9f5a-9d221dd6fd3fCited by top-tier papers5
- RL for Reasoning by Adaptively Revealing RationalesMohammad Hossein Amani, Aryo Lotfi, Nicolas Baldwin, Samy Bengio et al.ICLR 2026 · 19 citations
- On the Complexity of Learning Sparse Functions with Statistical and Gradient QueriesNirmit Joshi, Theodor Misiakiewicz, Nati SrebroNeurIPS 2024 · 16 citations
- Random Scaling of Emergent CapabilitiesRosie Zhao, Tian Qin, David Alvarez-Melis, Sham Kakade et al.ICML 2026 · 3 citations
- Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from -ParityJianhao Huang, Baharan MirzasoleimanICML 2026 · 2 citations
- Can Neural Networks Achieve Optimal Computational-statistical Tradeoff? An Analysis on Single-Index ModelSiyu Chen, Beining Wu, Miao Lu, Zhuoran Yang et al.ICLR 2025
Builds on17
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
Related papers
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade et al.NeurIPS 2022 · 220 citations
- How many degrees of freedom do we need to train deep networks: a loss landscape perspectiveBrett W. Larsen, Stanislav Fort, Nic Becker, Surya GanguliICLR 2022 · 33 citations
- Matching the Statistical Query Lower Bound for k-Sparse Parity Problems with Sign Stochastic Gradient DescentYiwen Kou, Zixiang Chen, Quanquan Gu, Sham M. KakadeNeurIPS 2024 · 7 citations
- Plant 'n' Seek: Can You Find the Winning Ticket?Jonas Fischer, Rebekka BurkholzICLR 2022 · 21 citations
- Sparse tree-based Initialization for Neural NetworksPatrick Lutz, Ludovic Arnould, Claire Boyer, Erwan ScornetICLR 2023
