A Non-Asymptotic Moreau Envelope Theory for High-Dimensional Generalized Linear Models
Lijia Zhou, Frederic Koehler, Pragya Sur, Danica J. Sutherland, Nati Srebro
Abstract
We prove a new generalization bound that shows for any class of linear predictors in Gaussian space, the Rademacher complexity of the class and the training error under any continuous loss can control the test error under all Moreau envelopes of the loss . We use our finite-sample bound to directly recover the"optimistic rate"of Zhou et al. (2021) for linear regression with the square loss, which is known to be tight for minimal -norm interpolation, but we also handle more general settings where the label is generated by a potentially misspecified multi-index model. The same argument can analyze noisy interpolation of max-margin classifiers through the squared hinge loss, and establishes consistency results in spiked-covariance settings. More generally, when the loss is only assumed to be Lipschitz, our bound effectively improves Talagrand's well-known contraction lemma by a factor of two, and we prove uniform convergence of interpolators (Koehler et al. 2021) for all smooth, non-negative losses. Finally, we show that application of our generalization bound using localized Gaussian width will generally be sharp for empirical risk minimizers, establishing a non-asymptotic Moreau envelope theory for generalization that applies outside of proportional scaling regimes, handles model misspecification, and complements existing asymptotic Moreau envelope theories for M-estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- H-Consistency Guarantees for RegressionAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2024 · 18 citations
- Noisy Interpolation Learning with Shallow Univariate ReLU NetworksNirmit Joshi, Gal Vardi, Nathan SrebroICLR 2024 · 12 citations
- Finite-Sample Analysis of Learning High-Dimensional Single ReLU NeuronJingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman et al.ICML 2023 · 9 citations
- Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or DimensionalityMarko Medvedev, Gal Vardi, Nati SrebroNeurIPS 2024 · 9 citations
- Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge RegressionAnvit Garg, Sohom Bhattacharya, Pragya SurICLR 2026 · 9 citations
Builds on2
- Uniform Convergence of Interpolators: Gaussian Width, Norm Bounds and Benign OverfittingFrederic Koehler, Lijia Zhou, Danica J. Sutherland, Nathan SrebroNeurIPS 2021 · 65 citations
- Multiple Descent: Design Your Own Generalization CurveLin Chen, Yifei Min, Mikhail Belkin, Amin KarbasiNeurIPS 2021 · 64 citations
Related papers
- Uniform Convergence with Square-Root Lipschitz LossLijia Zhou, Zhen Dai, Frederic Koehler, Nati SrebroNeurIPS 2023 · 2 citations
- On Contraction of Sequential and Offset Rademacher ComplexitiesAdam Block, Alexander Rakhlin, Mark SellkeICML 2026
- Tight and Fast Bounds for Multi-Label LearningYifan Zhang, Min-Ling ZhangICML 2025
- Optimistic Bounds for Multi-output LearningHenry W. J. Reeve, Ata KabánICML 2020 · 14 citations
- Generalization Analysis for Controllable LearningYifan Zhang, Xiao Zhang, Min-Ling ZhangICML 2025
