A theory of high dimensional regression with arbitrary correlations between input features and target functions: sample complexity, multiple descent curves and a hierarchy of phase transitions
Gabriel Mel, Surya Ganguli
摘要
The performance of neural networks depends on precise relationships between four distinct ingredients: the architecture, the loss function, the statistical structure of inputs, and the ground truth target function. Much theoretical work has focused on understanding the role of the first two ingredients under highly simplified models of random uncorrelated data and target functions. In contrast, performance likely relies on a conspiracy between the statistical structure of the input distribution and the structure of the function to be learned. To understand this better we revisit ridge regression in high dimensions, which corresponds to an exceedingly simple architecture and loss function, but we analyze its performance under arbitrary correlations between input features and the target function. We find a rich mathematical structure that includes: (1) a dramatic reduction in sample complexity when the target function aligns with data anisotropy; (2) the existence of multiple descent curves; (3) a sequence of phase transitions in the performance, loss landscape, and optimal regularization as a function of the amount of data that explains the first two effects.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations GeneralizeAlexander Wei, Wei Hu, Jacob SteinhardtICML 2022 · 被引用 90 次
- Overparameterization Improves Robustness to Covariate Shift in High DimensionsNilesh Tripuraneni, Ben Adlam, Jeffrey PenningtonNeurIPS 2021 · 被引用 50 次
- Optimal Ridge Regularization for Out-of-Distribution PredictionPratik Patil, Jin-Hong Du, Ryan J. TibshiraniICML 2024 · 被引用 23 次
- Benign Overfitting in Two-Layer ReLU Convolutional Neural Networks for XOR DataXuran Meng, Difan Zou, Yuan CaoICML 2024 · 被引用 11 次
- Information bottleneck theory of high-dimensional regression: relevancy, efficiency and optimalityVudtiwat Ngampruetikorn, David J. SchwabNeurIPS 2022 · 被引用 11 次
它引用的顶会 Paper1
相关 Paper
- Anisotropic Random Feature Regression in High DimensionsGabriel Mel, Jeffrey PenningtonICLR 2022 · 被引用 10 次
- Spurious Valleys and Clustering Behavior of Neural NetworksSamuele PollaciICML 2023 · 被引用 1 次
- Bayes-optimal Learning of Deep Random Networks of Extensive-widthHugo Cui, Florent Krzakala, Lenka ZdeborováICML 2023 · 被引用 49 次
- Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimizationBenjamin Aubin, Florent Krzakala, Yue M. Lu, Lenka ZdeborováNeurIPS 2020 · 被引用 67 次
- Learning Curves for SGD on Structured FeaturesBlake Bordelon, Cengiz PehlevanICLR 2022 · 被引用 29 次
