Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit
Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Eran Malach, Cyril Zhang
摘要
There is mounting evidence of emergent phenomena in the capabilities of deep learning methods as we scale up datasets, model sizes, and training times. While there are some accounts of how these resources modulate statistical capacity, far less is known about their effect on the computational problem of model training. This work conducts such an exploration through the lens of learning a -sparse parity of bits, a canonical discrete search problem which is statistically easy but computationally hard. Empirically, we find that a variety of neural networks successfully learn sparse parities, with discontinuous phase transitions in the training curves. On small instances, learning abruptly occurs at approximately iterations; this nearly matches SQ lower bounds, despite the apparent lack of a sparse prior. Our theoretical analysis shows that these observations are not explained by a Langevin-like mechanism, whereby SGD"stumbles in the dark"until it finds the hidden set of features (a natural algorithm which also runs in time). Instead, we show that SGD gradually amplifies the sparse solution via a Fourier gap in the population gradient, making continual progress that is invisible to loss and error metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper94
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 被引用 233 次
- The Clock and the Pizza: Two Stories in Mechanistic Explanation of Neural NetworksZiqian Zhong, Ziming Liu, Max Tegmark, Jacob AndreasNeurIPS 2023 · 被引用 181 次
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 被引用 179 次
- A Toy Model of Universality: Reverse Engineering how Networks Learn Group OperationsBilal Chughtai, Lawrence Chan, Neel NandaICML 2023 · 被引用 144 次
- The Evolution of Statistical Induction Heads: In-Context Learning Markov ChainsEzra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, Eran Malach 等NeurIPS 2024 · 被引用 140 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim 等NeurIPS 2020 · 被引用 397 次
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz 等NeurIPS 2022 · 被引用 267 次
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 被引用 193 次
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
相关 Paper
- Pareto Frontiers in Deep Feature Learning: Data, Compute, Width, and LuckBenjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Eran Malach 等NeurIPS 2023 · 被引用 8 次
- Matching the Statistical Query Lower Bound for k-Sparse Parity Problems with Sign Stochastic Gradient DescentYiwen Kou, Zixiang Chen, Quanquan Gu, Sham M. KakadeNeurIPS 2024 · 被引用 7 次
- An exactly solvable model for emergence and scaling laws in the multitask sparse parity problemYoonsoo Nam, Nayara Fonseca, Seok Hyeong Lee, Chris Mingard 等NeurIPS 2024 · 被引用 20 次
- On the universality of deep learningEmmanuel Abbe, Colin SandonNeurIPS 2020 · 被引用 29 次
- Feature learning via mean-field Langevin dynamics: classifying sparse parities and beyondTaiji Suzuki, Denny Wu, Kazusato Oko, Atsushi NitandaNeurIPS 2023 · 被引用 17 次
