Aggregation Hides Out-of-Distribution Generalization Failures from Spurious Correlations
Olawale Salaudeen, Haoran Zhang, Kumail Alhamoud, Sara Beery, Marzyeh Ghassemi
摘要
Benchmarks for out-of-distribution (OOD) generalization frequently show a strong positive correlation between in-distribution (ID) and OOD accuracy across models, termed "accuracy-on-the-line." This pattern is often taken to imply that spurious correlations-correlations that improve ID but reduce OOD performance-are rare in practice. We find that this positive correlation is often an artifact of aggregating heterogeneous OOD examples. Using a simple gradient-based method, OODSelect, we identify semantically coherent OOD subsets where accuracy on the line does not hold. Across widely used distribution shift benchmarks, the OODSelect uncovers subsets, sometimes up to over half of the standard OOD set, where higher ID accuracy predicts lower OOD accuracy. Our findings indicate that aggregate metrics can obscure important failure modes of OOD robustness. We release code and the identified subsets to facilitate further research. 2656 0.95 0.95 0.95 0.61 0.61 0.61 PACS Sketch 10 -0.48 (0.14) 0.37 (0.18) 0.17 (0.20) -0.05 (0.19) 0.39 (0.18) -0.06 (0.19) PACS Sketch 20 -0.33 (0.16) 0.71 (0.11) 0.17 (0.20) -0.41 (0.15) 0.34 (0.19) 0.00 (0.20) PACS Sketch 50 -0.47 (0.14) 0.84 (0.07) 0.00 (0.20) 0.23 (0.19) 0.47 (0.17) -0.12 (0.19) PACS Sketch 100 -0.33 (0.16) 0.73 (0.11) 0.11 (0.20) 0.29 (0.19) 0.70 (0.12) -0.08 (0.19) PACS Sketch 250 -0.30 (0.17) 0.79 (0.09) 0.19 (0.20) 0.35 (0.19) 0.70 (0.12) -0.04 (0.19) PACS Sketch 500 -0.22 (0.18) 0.83 (0.07) 0.28 (0.19) 0.41 (0.18) 0.69 (0.12) 0.07 (0.20) PACS Sketch 750 0.01 (0.20) 0.82 (0.08) 0.33 (0.19) 0.43 (0.17) 0.65 (0.13) 0.18 (0.20) PACS Sketch 800 0.05 (0.20) 0.81 (0.08) 0.33 (0.19) Dataset OOD N Pearson R Spearman ρ Ours Random Hard Ours Random Hard PACS Sketch 1500 0.21 (0.20) 0.82 (0.08) 0.41 (0.18) 0.42 (0.18) 0.68 (0.12) 0.39 (0.18) PACS Sketch 1750 0.24 (0.19) 0.82 (0.08) 0.46 (0.17) 0.42 (0.18) 0.67 (0.12) 0.43 (0.17) PACS Sketch 2000 0.29 (0.19) 0.82 (0.08) 0.49 (0.16) 0.48 (0.17) 0.66 (0.13) 0.48 (0.17) PACS Sketch 2250 0.48 (0.17) 0.81 (0.08) 0.54 (0.16) 0.52 (0.16) 0.67 (0.12) 0.51 (0.16) PACS Sketch 2500 0.56 (0.15) 0.81 (0.08) 0.58 (0.15) 0.54 (0.16) 0.67 (0.12) 0.56 (0.15) PACS Sketch 2750 0.62 (0.14) 0.81 (0.08) 0.63 (0.13) 0.56 (0.15) 0.67 (0.13) 0.61 (0.14) PACS Sketch 3000 0.67 (0.12) 0.81 (0.08) 0.68 (0.12) 0.58 (0.15) 0.67 (0.12) 0.64 (0.13) PACS Sketch 3250 0.71 (0.11) 0.81 (0.08) 0.72 (0.11) 0.61 (0.14) 0.67 (0.13) 0.66 (0.13) PACS Sketch 3500 0.75 (0.10) 0.81 (0.08) 0.76 (0.10) 0.63 (0.14) 0.66 (0.13) 0.66 (0.13) PACS Sketch 3750 0.81 (0.08) 0.81 (0.08) 0.79 (0.09) 0.67 (0.13) 0.66 (0.13) 0.67 (0.13) PACS Sketch 3929 0.81 0.81 0.81 0.67 0.67 0.67
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
相关 Paper
- OODRobustBench: a Benchmark and Large-Scale Analysis of Adversarial Robustness under Distribution ShiftLin Li, Yifei Wang, Chawin Sitawarin, Michael W. SpratlingICML 2024 · 被引用 13 次
- Assaying Out-Of-Distribution Generalization in Transfer LearningFlorian Wenzel, Andrea Dittadi, Peter V. Gehler, Carl-Johann Simon-Gabriel 等NeurIPS 2022 · 被引用 93 次
- ID and OOD Performance Are Sometimes Inversely Correlated on Real-world DatasetsDamien Teney, Yong Lin, Seong Joon Oh, Ehsan AbbasnejadNeurIPS 2023 · 被引用 70 次
- Agreement-on-the-line: Predicting the Performance of Neural Networks under Distribution ShiftChristina Baek, Yiding Jiang, Aditi Raghunathan, J. Zico KolterNeurIPS 2022 · 被引用 120 次
- OoD-Bench: Quantifying and Understanding Two Dimensions of Out-of-Distribution GeneralizationNanyang Ye, Kaican Li, Haoyue Bai, Runpeng Yu 等CVPR 2022 · 被引用 74 次
