Generalization Through the Lens of Leave-One-Out Error
Gregor Bachmann, Thomas Hofmann, Aurélien Lucchi
摘要
Despite the tremendous empirical success of deep learning models to solve various learning tasks, our theoretical understanding of their generalization ability is very limited. Classical generalization bounds based on tools such as the VC dimension or Rademacher complexity, are so far unsuitable for deep models and it is doubtful that these techniques can yield tight bounds even in the most idealistic settings (Nagarajan & Kolter, 2019) . In this work, we instead revisit the concept of leave-one-out (LOO) error to measure the generalization ability of deep models in the so-called kernel regime. While popular in statistics, the LOO error has been largely overlooked in the context of deep learning. By building upon the recently established connection between neural networks and kernel learning, we leverage the closed-form expression for the leave-one-out error, giving us access to an efficient proxy for the test error. We show both theoretically and empirically that the leave-one-out error is capable of capturing various phenomena in generalization theory, such as double descent, random labels or transfer learning. Our work therefore demonstrates that the leave-one-out error provides a tractable way to estimate the generalization ability of deep neural networks in the kernel regime, opening the door to potential, new research directions in the field of generalization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc 等ICML 2023 · 被引用 260 次
- A Fast, Well-Founded Approximation to the Empirical Neural Tangent KernelMohamad Amin Mohamadi, Wonho Bae, Danica J. SutherlandICML 2023 · 被引用 34 次
- The Memory-Perturbation Equation: Understanding Model's Sensitivity to DataPeter Nickl, Lu Xu, Dharmesh Tailor, Thomas Möllenhoff 等NeurIPS 2023 · 被引用 17 次
- Causal Estimation of Memorisation ProfilesPietro Lesci, Clara Meister, Thomas Hofmann, Andreas Vlachos 等ACL 2024 · 被引用 3 次
- Efficient Conditionally Invariant Representation LearningRoman Pogodin, Namrata Deka, Yazhe Li, Danica J. Sutherland 等ICLR 2023 · 被引用 2 次
它引用的顶会 Paper12
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee 等ICLR 2020 · 被引用 254 次
- Evaluation of Neural Architectures trained with square Loss vs Cross-Entropy in Classification TasksLike Hui, Mikhail BelkinICLR 2021 · 被引用 199 次
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 被引用 163 次
- Infinite attention: NNGP and NTK for deep attention networksJiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, Roman NovakICML 2020 · 被引用 147 次
相关 Paper
- Generalization Error of Generalized Linear Models in High DimensionsMelikasadat Emami, Mojtaba Sahraee-Ardakan, Parthe Pandit, Sundeep Rangan 等ICML 2020 · 被引用 40 次
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 被引用 133 次
- Nearly-tight Bounds for Deep Kernel LearningYifan Zhang, Min-Ling ZhangICML 2023 · 被引用 3 次
- On Leave-One-Out Conditional Mutual Information For GeneralizationMohamad Rida Rammal, Alessandro Achille, Aditya Golatkar, Suhas N. Diggavi 等NeurIPS 2022 · 被引用 11 次
- Phenomenology of Double Descent in Finite-Width Neural NetworksSidak Pal Singh, Aurélien Lucchi, Thomas Hofmann, Bernhard SchölkopfICLR 2022 · 被引用 12 次
