Approximate Cross-Validation for Structured Models
Soumya Ghosh, William T. Stephenson, Tin D. Nguyen, Sameer K. Deshpande, Tamara Broderick
摘要
Many modern data analyses benefit from explicitly modeling dependence structure in data -such as measurements across time or space, ordered words in a sentence, or genes in a genome. A gold standard evaluation technique is structured cross-validation (CV), which leaves out some data subset (such as data within a time interval or data in a geographic region) in each fold. But CV here can be prohibitively slow due to the need to re-run already-expensive learning algorithms many times. Previous work has shown approximate cross-validation (ACV) methods provide a fast and provably accurate alternative in the setting of empirical risk minimization. But this existing ACV work is restricted to simpler models by the assumptions that (i) data across CV folds are independent and (ii) an exact initial model fit is available. In structured data analyses, both these assumptions are often untrue. In the present work, we address (i) by extending ACV to CV schemes with dependence structure between the folds. To address (ii), we verify -both theoretically and empirically -that ACV quality deteriorates smoothly with noise in the initial fit. We demonstrate the accuracy and computational benefits of our proposed methods on a diverse set of real-world applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Cross-validation Confidence Intervals for Test ErrorPierre Bayle, Alexandre Bayle, Lucas Janson, Lester MackeyNeurIPS 2020 · 被引用 76 次
- Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without RefittingPrasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre L. Dognin 等NeurIPS 2022 · 被引用 36 次
- Iterative Approximate Cross-ValidationYuetian Luo, Zhimei Ren, Rina BarberICML 2023 · 被引用 9 次
- Synthetic data for model selectionAlon Shoshan, Nadav Bhonker, Igor Kviatkovsky, Matan Fintz 等ICML 2023 · 被引用 8 次
- Dropping Just a Handful of Preferences Can Change Top Large Language Model RankingsJenny Y. Huang, Yunyi Shen, Dennis Wei, Tamara BroderickICLR 2026 · 被引用 8 次
相关 Paper
- Approximate Cross-Validation with Low-Rank Data in High DimensionsWilliam T. Stephenson, Madeleine Udell, Tamara BroderickNeurIPS 2020 · 被引用 2 次
- General Approximate Cross Validation for Model Selection: Supervised, Semi-supervised and Pairwise LearningBowei Zhu, Yong LiuACM MM 2021 · 被引用 3 次
- Cross-Validation for Longitudinal Datasets with Unstable CorrelationsMeera Krishnamoorthy, Michael W. Sjoding, Jenna WiensKDD 2025
- Is Cross-validation the Gold Standard to Estimate Out-of-sample Model Performance?Garud Iyengar, Henry Lam, Tianyu WangNeurIPS 2024 · 被引用 6 次
- Ridge Regression: Structure, Cross-Validation, and SketchingSifan Liu, Edgar DobribanICLR 2020 · 被引用 52 次
