On the Optimal Weighted Regularization in Overparameterized Linear Regression
Denny Wu, Ji Xu
Abstract
We consider the linear model with in the overparameterized regime . We estimate via generalized (weighted) ridge regression: , where is the weighting matrix. Assuming a random effects model with general data covariance and anisotropic prior on the true coefficients , i.e., , we provide an exact characterization of the prediction risk in the proportional asymptotic limit . Our general setup leads to a number of interesting findings. We outline precise conditions that decide the sign of the optimal setting for the ridge parameter and confirm the implicit regularization effect of overparameterization, which theoretically justifies the surprising empirical observation that can be negative in the overparameterized regime. We also characterize the double descent phenomenon for principal component regression (PCR) when and are non-isotropic. Finally, we determine the optimal for both the ridgeless () and optimally regularized () case, and demonstrate the advantage of the weighted objective over standard ridge regression and PCR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- The Benefits of Implicit Regularization from SGD in Least Squares ProblemsDifan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu et al.NeurIPS 2021 · 41 citations
- Are Gaussian Data All You Need? The Extents and Limits of Universality in High-Dimensional Generalized Linear EstimationLuca Pesce, Florent Krzakala, Bruno Loureiro, Ludovic StephanICML 2023 · 6 citations
- Least Squares Regression Can Exhibit Under-Parameterized Double DescentXinyue Li, Rishi SonthaliaNeurIPS 2024 · 5 citations
- A Random Matrix Theory of Masked Self-Supervised LearningArie Zurich, Federica Gerace, Bruno Loureiro, Yue LuICML 2026
- On the Interplay between Graph Structure and Learning Algorithms in Graph Neural NetworksJunwei Su, Chuan WuICML 2025
Builds on4
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
- Exact expressions for double descent and implicit regularization via surrogate random designMichal Derezinski, Feynman T. Liang, Michael W. MahoneyNeurIPS 2020 · 81 citations
- Ridge Regression: Structure, Cross-Validation, and SketchingSifan Liu, Edgar DobribanICLR 2020 · 52 citations
- When does preconditioning help or hurt generalization?Shun-ichi Amari, Jimmy Ba, Roger Baker Grosse, Xuechen Li et al.ICLR 2021 · 11 citations
Related papers
- No Double Descent in Principal Component Regression: A High-Dimensional AnalysisDaniel Gedon, Antônio H. Ribeiro, Thomas B. SchönICML 2024 · 6 citations
- Anisotropic Random Feature Regression in High DimensionsGabriel Mel, Jeffrey PenningtonICLR 2022 · 10 citations
- On Optimal Interpolation in Linear RegressionEduard Oravkin, Patrick RebeschiniNeurIPS 2021 · 6 citations
- Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge RegressionAnvit Garg, Sohom Bhattacharya, Pragya SurICLR 2026 · 9 citations
- Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or DimensionalityMarko Medvedev, Gal Vardi, Nati SrebroNeurIPS 2024 · 9 citations
