On the non-universality of deep learning: quantifying the cost of symmetry
Emmanuel Abbe, Enric Boix-Adserà
摘要
We prove limitations on what neural networks trained by noisy gradient descent (GD) can efficiently learn. Our results apply whenever GD training is equivariant, which holds for many standard architectures and initializations. As applications, (i) we characterize the functions that fully-connected networks can weak-learn on the binary hypercube and unit sphere, demonstrating that depth-2 is as powerful as any other depth for this task; (ii) we extend the merged-staircase necessity result for learning with latent low-dimensional structure [ABM22] to beyond the mean-field regime. Under cryptographic assumptions, we also show hardness results for learning with fully-connected networks trained by stochastic gradient descent (SGD).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- How Far Can Transformers Reason? The Globality Barrier and Inductive ScratchpadEmmanuel Abbe, Samy Bengio, Aryo Lotfi, Colin Sandon 等NeurIPS 2024 · 被引用 52 次
- Learning in the Presence of Low-dimensional Structure: A Spiked Random Matrix PerspectiveJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2023 · 被引用 47 次
- Provable Advantage of Curriculum Learning on Parity Targets with Mixed InputsEmmanuel Abbe, Elisabetta Cornacchia, Aryo LotfiNeurIPS 2023 · 被引用 29 次
- When can transformers reason with abstract symbols?Enric Boix-Adserà, Omid Saremi, Emmanuel Abbe, Samy Bengio 等ICLR 2024 · 被引用 21 次
- Wait, Wait, Wait... Why Do Reasoning Models Loop?Charilaos Pipis, Shivam Garg, Vasilis Kontonis, Vaishnavi Shrivastava 等ICML 2026 · 被引用 17 次
它引用的顶会 Paper11
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- Lorentz Group Equivariant Neural Network for Particle PhysicsAlexander Bogatskiy, Brandon M. Anderson, Jan T. Offermann, Marwah Roussi 等ICML 2020 · 被引用 164 次
- Provably Strict Generalisation Benefit for Equivariant ModelsBryn Elesedy, Sheheryar ZaidiICML 2021 · 被引用 100 次
- The staircase property: How hierarchical structure can guide deep learningEmmanuel Abbe, Enric Boix-Adserà, Matthew S. Brennan, Guy Bresler 等NeurIPS 2021 · 被引用 74 次
- On the Sample Complexity of Learning under Geometric StabilityAlberto Bietti, Luca Venturi, Joan BrunaNeurIPS 2021 · 被引用 45 次
相关 Paper
- On the universality of deep learningEmmanuel Abbe, Colin SandonNeurIPS 2020 · 被引用 29 次
- Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label NoiseSpencer Frei, Yuan Cao, Quanquan GuICML 2021 · 被引用 22 次
- Computational Complexity of Learning Neural Networks: Smoothness and DegeneracyAmit Daniely, Nati Srebro, Gal VardiNeurIPS 2023 · 被引用 11 次
- Provable Guarantees for Neural Networks via Gradient Feature LearningZhenmei Shi, Junyi Wei, Yingyu LiangNeurIPS 2023 · 被引用 15 次
- Non-Singularity of the Gradient Descent Map for Neural Networks with Piecewise Analytic ActivationsAlexandru Craciun, Debarghya GhoshdastidarNeurIPS 2025 · 被引用 1 次
