Risk and cross validation in ridge regression with correlated samples
Alexander B. Atanasov, Jacob A. Zavatone-Veth, Cengiz Pehlevan
摘要
Recent years have seen substantial advances in our understanding of high-dimensional ridge regression, but existing theories assume that training examples are independent. By leveraging techniques from random matrix theory and free probability, we provide sharp asymptotics for the in-and out-of-sample risks of ridge regression when the data points have arbitrary correlations. We demonstrate that in this setting, the generalized cross validation estimator (GCV) fails to correctly predict the out-of-sample risk. However, in the case where the noise residuals have the same correlations as the data points, one can modify the GCV to yield an efficiently-computable unbiased estimator that concentrates in the high-dimensional limit, which we dub CorrGCV. We further extend our asymptotic analysis to the case where the test point has nontrivial correlations with the training set, a setting often encountered in time series forecasting. Assuming knowledge of the correlation structure of the time series, this again yields an extension of the GCV estimator, and sharply characterizes the degree to which such test points yield an overly optimistic prediction of long-time risk. We validate the predictions of our theory across a variety of high dimensional data. Statistics classically assumes that one has access to independent and identically distributed (i.i.d.) samples. However, this fundamental assumption is often violated when one considers data sampled from a time series-e.g., in the case of financial, climate, or neuroscience data (Bouchaud
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- On the Mechanisms of Weak-to-Strong Generalization: A Theoretical PerspectiveBehrad Moniri, Hamed HassaniNeurIPS 2025 · 被引用 8 次
- Pretrain–Test Task Alignment Governs Generalization in In-Context LearningMary Letey, Jacob A Zavatone-Veth, Yue M. Lu, Cengiz PehlevanICLR 2026 · 被引用 6 次
- A Random Matrix Perspective on the Consistency of Diffusion ModelsBinxu Wang, Jacob A Zavatone-Veth, Cengiz PehlevanICML 2026 · 被引用 4 次
- To Augment or Not to Augment? Diagnosing Distributional Symmetry BreakingHannah Lawrence, Elyssa F. Hofgard, Vasco Portilheiro, Yuxuan Chen 等ICLR 2026 · 被引用 1 次
它引用的顶会 Paper12
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 被引用 245 次
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard 等ICML 2020 · 被引用 184 次
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt 等NeurIPS 2021 · 被引用 170 次
- Understanding Double Descent Requires A Fine-Grained Bias-Variance DecompositionBen Adlam, Jeffrey PenningtonNeurIPS 2020 · 被引用 111 次
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy RegimeHugo Cui, Bruno Loureiro, Florent Krzakala, Lenka ZdeborováNeurIPS 2021 · 被引用 109 次
相关 Paper
- Asymptotically Free Sketched Ridge Ensembles: Risks, Cross-Validation, and TuningPratik Patil, Daniel LeJeuneICLR 2024 · 被引用 13 次
- Ridge Regression: Structure, Cross-Validation, and SketchingSifan Liu, Edgar DobribanICLR 2020 · 被引用 52 次
- Flat Minima in Linear Estimation and an Extended Gauss Markov TheoremSimon N. SegertICLR 2024 · 被引用 1 次
- A theory of high dimensional regression with arbitrary correlations between input features and target functions: sample complexity, multiple descent curves and a hierarchy of phase transitionsGabriel Mel, Surya GanguliICML 2021 · 被引用 24 次
- Subsample Ridge Ensembles: Equivalences and Generalized Cross-ValidationJin-Hong Du, Pratik Patil, Arun K. KuchibhotlaICML 2023 · 被引用 12 次
