Optimal Ridge Regularization for Out-of-Distribution Prediction
Pratik Patil, Jin-Hong Du, Ryan J. Tibshirani
摘要
We study the behavior of optimal ridge regularization and optimal ridge risk for out-of-distribution prediction, where the test distribution deviates arbitrarily from the train distribution. We establish general conditions that determine the sign of the optimal regularization level under covariate and regression shifts. These conditions capture the alignment between the covariance and signal structures in the train and test data and reveal stark differences compared to the in-distribution setting. For example, a negative regularization level can be optimal under covariate shift or regression shift, even when the training features are isotropic or the design is underparameterized. Furthermore, we prove that the optimally-tuned risk is monotonic in the data aspect ratio, even in the out-of-distribution setting and when optimizing over negative regularization levels. In general, our results do not make any modeling assumptions for the train or the test distributions, except for moment bounds, and allow for arbitrary shifts and the widest possible range of (negative) regularization levels.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge RegressionAnvit Garg, Sohom Bhattacharya, Pragya SurICLR 2026 · 被引用 9 次
- Pretrain–Test Task Alignment Governs Generalization in In-Context LearningMary Letey, Jacob A Zavatone-Veth, Yue M. Lu, Cengiz PehlevanICLR 2026 · 被引用 6 次
- High-dimensional Analysis of Synthetic Data SelectionParham Rezaei, Filip Kovacevic, Francesco Locatello, Marco MondelliICLR 2026 · 被引用 6 次
- Generalization vs Specialization under Concept ShiftAlex Nguyen, David J. Schwab, Vudtiwat NgampruetikornNeurIPS 2025 · 被引用 3 次
- Transfer Learning for Benign Overfitting in High-Dimensional Linear RegressionYeichan Kim, Ilmun Kim, Seyoung ParkNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper11
- Rethinking Bias-Variance Trade-off for Generalization of Neural NetworksZitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt 等ICML 2020 · 被引用 219 次
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
- Optimal Regularization can Mitigate Double DescentPreetum Nakkiran, Prayaag Venkat, Sham M. Kakade, Tengyu MaICLR 2021 · 被引用 148 次
- More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations GeneralizeAlexander Wei, Wei Hu, Jacob SteinhardtICML 2022 · 被引用 90 次
- Kernel Alignment Risk Estimator: Risk Prediction from Training DataArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler 等NeurIPS 2020 · 被引用 74 次
相关 Paper
- Generalized equivalences between subsampling and ridge regularizationPratik Patil, Jin-Hong DuNeurIPS 2023 · 被引用 10 次
- Overparameterization Improves Robustness to Covariate Shift in High DimensionsNilesh Tripuraneni, Ben Adlam, Jeffrey PenningtonNeurIPS 2021 · 被引用 50 次
- Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot TuningHien Dang, Pratik Patil, Alessandro RinaldoICML 2026 · 被引用 1 次
- Benign Overfitting in Out-of-Distribution Generalization of Linear ModelsShange Tang, Jiayun Wu, Jianqing Fan, Chi JinICLR 2025
- Optimal Regularization for Performative LearningEdwige Cyffers, Alireza Mirrokni, Marco MondelliICML 2026 · 被引用 1 次
