Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
Olawale Salaudeen, Haoran Zhang, Kumail Alhamoud, Sara Beery, Marzyeh Ghassemi
Abstract
Benchmarks for out-of-distribution (OOD) generalization frequently show a strong positive correlation between in-distribution (ID) and OOD accuracy across models, termed "accuracy-on-the-line." This pattern is often taken to imply that spurious correlations-correlations that improve ID but reduce OOD performance-are rare in practice. We find that this positive correlation is often an artifact of aggregating heterogeneous OOD examples. Using a simple gradient-based method, OODSelect, we identify semantically coherent OOD subsets where accuracy on the line does not hold. Across widely used distribution shift benchmarks, the OODSelect uncovers subsets, sometimes up to over half of the standard OOD set, where higher ID accuracy predicts lower OOD accuracy. Our findings indicate that aggregate metrics can obscure important failure modes of OOD robustness. We release code and the identified subsets to facilitate further research. 2656 0.95 0.95 0.95 0.61 0.61 0.61 PACS Sketch 10 -0.48 (0.14) 0.37 (0.18) 0.17 (0.20) -0.05 (0.19) 0.39 (0.18) -0.06 (0.19) PACS Sketch 20 -0.33 (0.16) 0.71 (0.11) 0.17 (0.20) -0.41 (0.15) 0.34 (0.19) 0.00 (0.20) PACS Sketch 50 -0.47 (0.14) 0.84 (0.07) 0.00 (0.20) 0.23 (0.19) 0.47 (0.17) -0.12 (0.19) PACS Sketch 100 -0.33 (0.16) 0.73 (0.11) 0.11 (0.20) 0.29 (0.19) 0.70 (0.12) -0.08 (0.19) PACS Sketch 250 -0.30 (0.17) 0.79 (0.09) 0.19 (0.20) 0.35 (0.19) 0.70 (0.12) -0.04 (0.19) PACS Sketch 500 -0.22 (0.18) 0.83 (0.07) 0.28 (0.19) 0.41 (0.18) 0.69 (0.12) 0.07 (0.20) PACS Sketch 750 0.01 (0.20) 0.82 (0.08) 0.33 (0.19) 0.43 (0.17) 0.65 (0.13) 0.18 (0.20) PACS Sketch 800 0.05 (0.20) 0.81 (0.08) 0.33 (0.19) Dataset OOD N Pearson R Spearman ρ Ours Random Hard Ours Random Hard PACS Sketch 1500 0.21 (0.20) 0.82 (0.08) 0.41 (0.18) 0.42 (0.18) 0.68 (0.12) 0.39 (0.18) PACS Sketch 1750 0.24 (0.19) 0.82 (0.08) 0.46 (0.17) 0.42 (0.18) 0.67 (0.12) 0.43 (0.17) PACS Sketch 2000 0.29 (0.19) 0.82 (0.08) 0.49 (0.16) 0.48 (0.17) 0.66 (0.13) 0.48 (0.17) PACS Sketch 2250 0.48 (0.17) 0.81 (0.08) 0.54 (0.16) 0.52 (0.16) 0.67 (0.12) 0.51 (0.16) PACS Sketch 2500 0.56 (0.15) 0.81 (0.08) 0.58 (0.15) 0.54 (0.16) 0.67 (0.12) 0.56 (0.15) PACS Sketch 2750 0.62 (0.14) 0.81 (0.08) 0.63 (0.13) 0.56 (0.15) 0.67 (0.13) 0.61 (0.14) PACS Sketch 3000 0.67 (0.12) 0.81 (0.08) 0.68 (0.12) 0.58 (0.15) 0.67 (0.12) 0.64 (0.13) PACS Sketch 3250 0.71 (0.11) 0.81 (0.08) 0.72 (0.11) 0.61 (0.14) 0.67 (0.13) 0.66 (0.13) PACS Sketch 3500 0.75 (0.10) 0.81 (0.08) 0.76 (0.10) 0.63 (0.14) 0.66 (0.13) 0.66 (0.13) PACS Sketch 3750 0.81 (0.08) 0.81 (0.08) 0.79 (0.09) 0.67 (0.13) 0.66 (0.13) 0.67 (0.13) PACS Sketch 3929 0.81 0.81 0.81 0.67 0.67 0.67
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
Related papers
- OODRobustBench: a Benchmark and Large-Scale Analysis of Adversarial Robustness under Distribution ShiftLin Li, Yifei Wang, Chawin Sitawarin, Michael W. SpratlingICML 2024 · 13 citations
- Assaying Out-Of-Distribution Generalization in Transfer LearningFlorian Wenzel, Andrea Dittadi, Peter V. Gehler, Carl-Johann Simon-Gabriel et al.NeurIPS 2022 · 93 citations
- ID and OOD Performance Are Sometimes Inversely Correlated on Real-world DatasetsDamien Teney, Yong Lin, Seong Joon Oh, Ehsan AbbasnejadNeurIPS 2023 · 70 citations
- Agreement-on-the-line: Predicting the Performance of Neural Networks under Distribution ShiftChristina Baek, Yiding Jiang, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 120 citations
- OoD-Bench: Quantifying and Understanding Two Dimensions of Out-of-Distribution GeneralizationNanyang Ye, Kaican Li, Haoyue Bai, Runpeng Yu et al.CVPR 2022 · 74 citations
