A theory of high dimensional regression with arbitrary correlations between input features and target functions: sample complexity, multiple descent curves and a hierarchy of phase transitions
Gabriel Mel, Surya Ganguli
Abstract
The performance of neural networks depends on precise relationships between four distinct ingredients: the architecture, the loss function, the statistical structure of inputs, and the ground truth target function. Much theoretical work has focused on understanding the role of the first two ingredients under highly simplified models of random uncorrelated data and target functions. In contrast, performance likely relies on a conspiracy between the statistical structure of the input distribution and the structure of the function to be learned. To understand this better we revisit ridge regression in high dimensions, which corresponds to an exceedingly simple architecture and loss function, but we analyze its performance under arbitrary correlations between input features and the target function. We find a rich mathematical structure that includes: (1) a dramatic reduction in sample complexity when the target function aligns with data anisotropy; (2) the existence of multiple descent curves; (3) a sequence of phase transitions in the performance, loss landscape, and optimal regularization as a function of the amount of data that explains the first two effects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28b216f2-70d7-4f89-8b31-710f199a4bc4Cited by top-tier papers13
- More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations GeneralizeAlexander Wei, Wei Hu, Jacob SteinhardtICML 2022 · 90 citations
- Overparameterization Improves Robustness to Covariate Shift in High DimensionsNilesh Tripuraneni, Ben Adlam, Jeffrey PenningtonNeurIPS 2021 · 50 citations
- Optimal Ridge Regularization for Out-of-Distribution PredictionPratik Patil, Jin-Hong Du, Ryan J. TibshiraniICML 2024 · 23 citations
- Benign Overfitting in Two-Layer ReLU Convolutional Neural Networks for XOR DataXuran Meng, Difan Zou, Yuan CaoICML 2024 · 11 citations
- Information bottleneck theory of high-dimensional regression: relevancy, efficiency and optimalityVudtiwat Ngampruetikorn, David J. SchwabNeurIPS 2022 · 11 citations
Builds on1
Related papers
- Anisotropic Random Feature Regression in High DimensionsGabriel Mel, Jeffrey PenningtonICLR 2022 · 10 citations
- Spurious Valleys and Clustering Behavior of Neural NetworksSamuele PollaciICML 2023 · 1 citation
- Bayes-optimal Learning of Deep Random Networks of Extensive-widthHugo Cui, Florent Krzakala, Lenka ZdeborováICML 2023 · 49 citations
- Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimizationBenjamin Aubin, Florent Krzakala, Yue M. Lu, Lenka ZdeborováNeurIPS 2020 · 67 citations
- Learning Curves for SGD on Structured FeaturesBlake Bordelon, Cengiz PehlevanICLR 2022 · 29 citations
