LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models
Fanfei Li, Thomas Klein, Wieland Brendel, Robert Geirhos, Roland S. Zimmermann
Abstract
Out-of-distribution (OOD) robustness is a desired property of computer vision models. Improving model robustness requires high-quality signals from robustness benchmarks to quantify progress. While various benchmark datasets such as ImageNet-C were proposed in the Im-ageNet era, most ImageNet-C corruption types are no longer OOD relative to today's large, webscraped datasets, which already contain common corruptions such as blur or JPEG compression artifacts. Consequently, these benchmarks are no longer well-suited for evaluating OOD robustness in the era of web-scale datasets. Indeed, recent models show saturating scores on ImageNet-era OOD benchmarks, indicating that it is unclear whether models trained on web-scale datasets truly become better at OOD generalization or whether they have simply been exposed to the test distortions during training. To address this, we introduce LAION-C as a benchmark alternative for ImageNet-C. LAION-C consists of six novel distortion types specifically designed to be OOD, even for web-scale datasets such as LAION. In a comprehensive evaluation of stateof-the-art models, we find that the LAION-C dataset poses significant challenges to contemporary models, including MLLMs such as Gemini and GPT-4o. We additionally conducted a psychophysical experiment to evaluate the difficulty of our corruptions for human observers, enabling a comparison of models to lab-quality human robustness data. We observe a paradigm shift in OOD generalization: from humans outperforming models, to the best models now matching or outperforming the best human observers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc76e113-bb76-4177-9c93-d63483f774d2Cited by top-tier papers5
- Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of ClassifiersThomas Klein, Sascha Meyen, Wieland Brendel, Felix A. Wichmann et al.NeurIPS 2025 · 2 citations
- Low-Pass Filtering Improves Behavioral Alignment of Vision ModelsMax Wolff, Thomas Klein, Evgenia Rusak, Felix A. Wichmann et al.ICLR 2026 · 2 citations
- Back to Source: Open-Set Continual Test-Time Adaptation via Domain CompensationYingkai Yang, Chaoqi Chen, Hui HuangCVPR 2026 · 1 citation
- Flow Matching for Multimodal DistributionsGaoxiang Luo, Frank Cole, Sihang Zhang, Yuxiang Wan et al.CVPR 2026 · 1 citation
- Self-Soupervision: Cooking Model Soups without LabelsAnthony Fuller, James Green, Evan ShelhamerICML 2026
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
Related papers
- In Search of Forgotten Domain GeneralizationPrasanna Mayilvahanan, Roland S. Zimmermann, Thaddäus Wiedemer, Evgenia Rusak et al.ICLR 2025
- CNS-Bench: Benchmarking Image Classifier Robustness Under Continuous Nuisance ShiftsOlaf Dünkel, Artur Jesslen, Jiahao Xie, Christian Theobalt et al.ICCV 2025
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution ShiftsJiansheng Li, Xingxuan Zhang, Hao Zou, Yige Guo et al.CVPR 2025
- Does CLIP's generalization performance mainly stem from high train-test similarity?Prasanna Mayilvahanan, Thaddäus Wiedemer, Evgenia Rusak, Matthias Bethge et al.ICLR 2024 · 43 citations
- Improving robustness against common corruptions by covariate shift adaptationSteffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann et al.NeurIPS 2020 · 688 citations
