The Relative Instability of Model Comparison with Cross-validation
Alexandre Bayle, Lucas Janson, Lester Mackey
Abstract
Cross-validation (CV) is known to provide asymptotically exact tests and confidence intervals for model improvement but only when the model comparison is relatively stable. Surprisingly, we prove that even simple, individually stable models can generate relatively unstable comparisons, calling into question the validity of CV inference. Specifically, we show that the Lasso and its close cousin, soft-thresholding, generate relatively unstable comparisons and invalid CV inferences, even in the most favorable of learning settings and even when both models are individually stable. These findings highlight the importance of verifying relative stability before deploying CV for model comparison.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e074db87-1338-4993-8ae8-53c17f19c1ecRelated papers
- Cross-validation Confidence Intervals for Test ErrorPierre Bayle, Alexandre Bayle, Lucas Janson, Lester MackeyNeurIPS 2020 · 76 citations
- Can we globally optimize cross-validation loss? Quasiconvexity in ridge regressionWilliam T. Stephenson, Zachary Frangella, Madeleine Udell, Tamara BroderickNeurIPS 2021 · 15 citations
- Deciphering Lasso-based Classification Through a Large Dimensional Analysis of the Iterative Soft-Thresholding AlgorithmMalik Tiomoko, Ekkehard Schnoor, Mohamed El Amine Seddik, Igor Colin et al.ICML 2022 · 4 citations
- Cross-Validation for Longitudinal Datasets with Unstable CorrelationsMeera Krishnamoorthy, Michael W. Sjoding, Jenna WiensKDD 2025
- Is Cross-validation the Gold Standard to Estimate Out-of-sample Model Performance?Garud Iyengar, Henry Lam, Tianyu WangNeurIPS 2024 · 6 citations
