Agreement-on-the-line: Predicting the Performance of Neural Networks under Distribution Shift
Christina Baek, Yiding Jiang, Aditi Raghunathan, J. Zico Kolter
Abstract
Recently , Miller et al. [59] showed that a model's in-distribution (ID) accuracy has a strong linear correlation with its out-of-distribution (OOD) accuracy on several OOD benchmarks -a phenomenon they dubbed "accuracy-on-the-line". While a useful tool for model selection (i.e., the models with better ID accuracy are likely to have better OOD accuracy), this fact does not help estimate the actual OOD performance of models without access to a labeled OOD validation set. In this paper, we show a similar but surprising phenomenon also holds for the agreement between pairs of neural network classifiers: whenever accuracy-onthe-line holds, we observe that the OOD agreement between the predictions of any two pairs of neural networks (with potentially different architectures) also observes a strong linear correlation with their ID agreement. Furthermore, we observe that the slope and bias of OOD vs. ID agreement closely matches that of OOD vs. ID accuracy. This phenomenon, which we call agreement-on-the-line, has important practical applications: without any labeled data, we can predict the OOD accuracy of classifiers, since OOD agreement can be estimated with just unlabeled data. Our prediction algorithm outperforms previous methods both in shifts where agreement-on-the-line holds and, surprisingly, when accuracy is not on the line. This phenomenon also provides new insights into deep neural networks: unlike accuracy-on-the-line, agreement-on-the-line appears to only hold for neural network classifiers. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b58092b-e047-4cad-9186-a75ab1888983Cited by top-tier papers45
- Assaying Out-Of-Distribution Generalization in Transfer LearningFlorian Wenzel, Andrea Dittadi, Peter V. Gehler, Carl-Johann Simon-Gabriel et al.NeurIPS 2022 · 93 citations
- Towards Last-layer Retraining for Group Robustness with Fewer AnnotationsTyler LaBonte, Vidya Muthukumar, Abhishek KumarNeurIPS 2023 · 73 citations
- On the Strong Correlation Between Model Invariance and GeneralizationWeijian Deng, Stephen Gould, Liang ZhengNeurIPS 2022 · 28 citations
- The Entropy Enigma: Success and Failure of Entropy MinimizationOri Press, Ravid Shwartz-Ziv, Yann LeCun, Matthias BethgeICML 2024 · 27 citations
- Certified Robustness via Dynamic Margin Maximization and Improved Lipschitz RegularizationMahyar Fazlyab, Taha Entesari, Aniket Roy, Rama ChellappaNeurIPS 2023 · 26 citations
Builds on28
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image ClassificationChun-Fu (Richard) Chen, Quanfu Fan, Rameswar PandaICCV 2021 · 2,072 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
Related papers
- Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-LineEungyeup Kim, Mingjie Sun, Christina Baek, Aditi Raghunathan et al.NeurIPS 2024 · 12 citations
- Predicting the Performance of Foundation Models via Agreement-on-the-LineRahul Saxena, Taeyoun Kim, Aman Mehra, Christina Baek et al.NeurIPS 2024 · 8 citations
- Demystifying Disagreement-on-the-Line in High DimensionsDonghwan Lee, Behrad Moniri, Xinmeng Huang, Edgar Dobriban et al.ICML 2023 · 12 citations
- Aggregation Hides Out-of-Distribution Generalization Failures from Spurious CorrelationsOlawale Salaudeen, Haoran Zhang, Kumail Alhamoud, Sara Beery et al.NeurIPS 2025 · 3 citations
- (Almost) Provable Error Bounds Under Distribution Shift via Disagreement DiscrepancyElan Rosenfeld, Saurabh GargNeurIPS 2023 · 18 citations
