On the Optimal Weighted Regularization in Overparameterized Linear Regression
Denny Wu, Ji Xu
摘要
We consider the linear model with in the overparameterized regime . We estimate via generalized (weighted) ridge regression: , where is the weighting matrix. Assuming a random effects model with general data covariance and anisotropic prior on the true coefficients , i.e., , we provide an exact characterization of the prediction risk in the proportional asymptotic limit . Our general setup leads to a number of interesting findings. We outline precise conditions that decide the sign of the optimal setting for the ridge parameter and confirm the implicit regularization effect of overparameterization, which theoretically justifies the surprising empirical observation that can be negative in the overparameterized regime. We also characterize the double descent phenomenon for principal component regression (PCR) when and are non-isotropic. Finally, we determine the optimal for both the ridgeless () and optimally regularized () case, and demonstrate the advantage of the weighted objective over standard ridge regression and PCR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- The Benefits of Implicit Regularization from SGD in Least Squares ProblemsDifan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu 等NeurIPS 2021 · 被引用 41 次
- Are Gaussian Data All You Need? The Extents and Limits of Universality in High-Dimensional Generalized Linear EstimationLuca Pesce, Florent Krzakala, Bruno Loureiro, Ludovic StephanICML 2023 · 被引用 6 次
- Least Squares Regression Can Exhibit Under-Parameterized Double DescentXinyue Li, Rishi SonthaliaNeurIPS 2024 · 被引用 5 次
- A Random Matrix Theory of Masked Self-Supervised LearningArie Zurich, Federica Gerace, Bruno Loureiro, Yue LuICML 2026
- On the Interplay between Graph Structure and Learning Algorithms in Graph Neural NetworksJunwei Su, Chuan WuICML 2025
它引用的顶会 Paper4
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 被引用 163 次
- Exact expressions for double descent and implicit regularization via surrogate random designMichal Derezinski, Feynman T. Liang, Michael W. MahoneyNeurIPS 2020 · 被引用 81 次
- Ridge Regression: Structure, Cross-Validation, and SketchingSifan Liu, Edgar DobribanICLR 2020 · 被引用 52 次
- When does preconditioning help or hurt generalization?Shun-ichi Amari, Jimmy Ba, Roger Baker Grosse, Xuechen Li 等ICLR 2021 · 被引用 11 次
相关 Paper
- No Double Descent in Principal Component Regression: A High-Dimensional AnalysisDaniel Gedon, Antônio H. Ribeiro, Thomas B. SchönICML 2024 · 被引用 6 次
- Anisotropic Random Feature Regression in High DimensionsGabriel Mel, Jeffrey PenningtonICLR 2022 · 被引用 10 次
- On Optimal Interpolation in Linear RegressionEduard Oravkin, Patrick RebeschiniNeurIPS 2021 · 被引用 6 次
- Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge RegressionAnvit Garg, Sohom Bhattacharya, Pragya SurICLR 2026 · 被引用 9 次
- Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or DimensionalityMarko Medvedev, Gal Vardi, Nati SrebroNeurIPS 2024 · 被引用 9 次
