Learning Curves for Deep Structured Gaussian Feature Models
Jacob A. Zavatone-Veth, Cengiz Pehlevan
摘要
In recent years, significant attention in deep learning theory has been devoted to analyzing when models that interpolate their training data can still generalize well to unseen examples. Many insights have been gained from studying models with multiple layers of Gaussian random features, for which one can compute precise generalization asymptotics. However, few works have considered the effect of weight anisotropy; most assume that the random features are generated using independent and identically distributed Gaussian weights, and allow only for structure in the input data. Here, we use the replica trick from statistical physics to derive learning curves for models with many layers of structured Gaussian features. We show that allowing correlations between the rows of the first layer of features can aid generalization, while structure in later layers is generally detrimental. Our results shed light on how weight structure affects generalization in a simple class of solvable models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- A Dynamical Model of Neural Scaling LawsBlake Bordelon, Alexander B. Atanasov, Cengiz PehlevanICML 2024 · 被引用 84 次
- Asymptotics of feature learning in two-layer networks after one gradient-stepHugo Cui, Luca Pesce, Yatin Dandi, Florent Krzakala 等ICML 2024 · 被引用 30 次
- Asymptotics of Learning with Deep Structured (Random) FeaturesDominik Schröder, Daniil Dmitriev, Hugo Cui, Bruno LoureiroICML 2024 · 被引用 12 次
- More is Better: when Infinite Overparameterization is Optimal and Overfitting is ObligatoryJames B. Simon, Dhruva Karkada, Nikhil Ghosh, Mikhail BelkinICLR 2024 · 被引用 7 次
- Demystifying LLM-as-a-Judge: Analytically Tractable Model for Inference-Time ScalingIndranil Halder, Cengiz PehlevanICML 2026 · 被引用 4 次
它引用的顶会 Paper16
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 被引用 245 次
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard 等ICML 2020 · 被引用 184 次
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
相关 Paper
- Anisotropic Random Feature Regression in High DimensionsGabriel Mel, Jeffrey PenningtonICLR 2022 · 被引用 10 次
- On the geometry of generalization and memorization in deep neural networksCory Stephenson, Suchismita Padhy, Abhinav Ganesh, Yue Hui 等ICLR 2021 · 被引用 95 次
- On the Inherent Regularization Effects of Noise Injection During TrainingOussama Dhifallah, Yue M. LuICML 2021 · 被引用 36 次
- On the interplay between data structure and loss function in classification problemsStéphane d'Ascoli, Marylou Gabrié, Levent Sagun, Giulio BiroliNeurIPS 2021 · 被引用 17 次
- Model, sample, and epoch-wise descents: exact solution of gradient flow in the random feature modelAntoine Bodin, Nicolas MacrisNeurIPS 2021 · 被引用 19 次
