Assessing Generalization of SGD via Disagreement
Yiding Jiang, Vaishnavh Nagarajan, Christina Baek, J. Zico Kolter
Abstract
We empirically show that the test error of deep networks can be estimated by simply training the same architecture on the same training set but with a different run of Stochastic Gradient Descent (SGD), and measuring the disagreement rate between the two networks on unlabeled test data. This builds on -- and is a stronger version of -- the observation in Nakkiran&Bansal '20, which requires the second run to be on an altogether fresh training set. We further theoretically show that this peculiar phenomenon arises from the well-calibrated nature of ensembles of SGD-trained models. This finding not only provides a simple empirical measure to directly predict the test error using unlabeled test data, but also establishes a new conceptual connection between generalization and calibration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8fcffac-0268-4679-8c58-df720d45ca59Cited by top-tier papers54
- Leveraging unlabeled data to predict out-of-distribution performanceSaurabh Garg, Sivaraman Balakrishnan, Zachary Chase Lipton, Behnam Neyshabur et al.ICLR 2022 · 160 citations
- Agreement-on-the-line: Predicting the Performance of Neural Networks under Distribution ShiftChristina Baek, Yiding Jiang, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 120 citations
- Detecting Errors and Estimating Accuracy on Unlabeled Data with Self-training EnsemblesJiefeng Chen, Frederick Liu, Besim Avci, Xi Wu et al.NeurIPS 2021 · 79 citations
- Towards Last-layer Retraining for Group Robustness with Fewer AnnotationsTyler LaBonte, Vidya Muthukumar, Abhishek KumarNeurIPS 2023 · 73 citations
- Datamodels: Understanding Predictions with Data and Data with PredictionsAndrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc et al.ICML 2022 · 66 citations
Builds on12
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
- Distribution-free binary classification: prediction sets, confidence intervals and calibrationChirag Gupta, Aleksandr Podkopaev, Aaditya RamdasNeurIPS 2020 · 105 citations
- Detecting Errors and Estimating Accuracy on Unlabeled Data with Self-training EnsemblesJiefeng Chen, Frederick Liu, Besim Avci, Xi Wu et al.NeurIPS 2021 · 79 citations
Related papers
- (Almost) Provable Error Bounds Under Distribution Shift via Disagreement DiscrepancyElan Rosenfeld, Saurabh GargNeurIPS 2023 · 18 citations
- On Predicting Generalization using GANsYi Zhang, Arushi Gupta, Nikunj Saunshi, Sanjeev AroraICLR 2022 · 8 citations
- Inconsistency, Instability, and Generalization Gap of Deep Neural Network TrainingRie Johnson, Tong ZhangNeurIPS 2023 · 11 citations
- Bad Global Minima Exist and SGD Can Reach ThemShengchao Liu, Dimitris S. Papailiopoulos, Dimitris AchlioptasNeurIPS 2020 · 89 citations
- Extreme Memorization via Scale of InitializationHarsh Mehta, Ashok Cutkosky, Behnam NeyshaburICLR 2021 · 22 citations
