Sketched Ridgeless Linear Regression: The Role of Downsampling
Xin Chen, Yicheng Zeng, Siyue Yang, Qiang Sun
Abstract
Overparametrization often helps improve the generalization performance. This paper presents a dual view of overparametrization suggesting that downsampling may also help generalize. Focusing on the proportional regime , where represents the sketching size, is the sample size, and is the feature dimensionality, we investigate two out-of-sample prediction risks of the sketched ridgeless least square estimator. Our findings challenge conventional beliefs by showing that downsampling does not always harm generalization but can actually improve it in certain cases. We identify the optimal sketching size that minimizes out-of-sample prediction risks and demonstrate that the optimally sketched estimator exhibits stabler risk curves, eliminating the peaks of those for the full-sample estimator. To facilitate practical implementation, we propose an empirical procedure to determine the optimal sketching size. Finally, we extend our analysis to cover central limit theorems and misspecified models. Numerical studies strongly support our theory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Asymptotically Free Sketched Ridge Ensembles: Risks, Cross-Validation, and TuningPratik Patil, Daniel LeJeuneICLR 2024 · 13 citations
- High-Dimensional Analysis for Generalized Nonlinear Regression: From Asymptotics to AlgorithmJian Li, Yong Liu, Weiping WangAAAI 2024 · 4 citations
- Implicit Regularization Paths of Weighted Neural RepresentationsJin-Hong Du, Pratik PatilNeurIPS 2024 · 2 citations
- Why Self-Distillation Helps and Hurts: Denoising vs. Signal ForgettingMingqi Wu, Archer Yang, Qiang SunICML 2026
- Prediction Risk and Estimation Risk of the Ridgeless Least Squares Estimator under General Assumptions on Regression ErrorsSungyoon Lee, Sokbae LeeICLR 2025
Builds on3
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Generalization of Two-layer Neural Networks: An Asymptotic ViewpointJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Denny Wu et al.ICLR 2020 · 77 citations
- Asymptotic Normality and Confidence Intervals for Prediction Risk of the Min-Norm Least Squares EstimatorZeng Li, Chuanlong Xie, Qinwen WangICML 2021 · 4 citations
Related papers
- Subsample Ridge Ensembles: Equivalences and Generalized Cross-ValidationJin-Hong Du, Pratik Patil, Arun K. KuchibhotlaICML 2023 · 12 citations
- A Fast and Accurate Estimator for Large Scale Linear Model via Data AveragingRui Wang, Yanyan Ouyang, Panpan Yu, Wangli XuNeurIPS 2023 · 1 citation
- Generalized equivalences between subsampling and ridge regularizationPratik Patil, Jin-Hong DuNeurIPS 2023 · 10 citations
- Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or DimensionalityMarko Medvedev, Gal Vardi, Nati SrebroNeurIPS 2024 · 9 citations
- No Free Lunch from Random Feature Ensembles: Scaling Laws and Near-Optimality ConditionsBenjamin S. Ruben, William Lingxiao Tong, Hamza Tahir Chaudhry, Cengiz PehlevanICML 2025
