ImageNet-X: Understanding Model Mistakes with Factor of Variation Annotations
Badr Youbi Idrissi, Diane Bouchacourt, Randall Balestriero, Ivan Evtimov, Caner Hazirbas, Nicolas Ballas, Pascal Vincent, Michal Drozdzal, David Lopez-Paz, Mark Ibrahim
Abstract
Deep learning vision systems are widely deployed across applications where reliability is critical. However, even today's best models can fail to recognize an object when its pose, lighting, or background varies. While existing benchmarks surface examples that are challenging for models, they do not explain why such mistakes arise. To address this need, we introduce ImageNet-X-a set of sixteen human annotations of factors such as pose, background, or lighting for the entire ImageNet-1k validation set as well as a random subset of 12k training images. Equipped with ImageNet-X, we investigate 2,200 current recognition models and study the types of mistakes as a function of model's (1) architecture -e.g. transformer vs. convolutional -, (2) learning paradigm -e.g. supervised vs. self-supervised -, and (3) training procedurese.g. data augmentation. Regardless of these choices, we find models have consistent failure modes across ImageNet-X categories. We also find that while data augmentation can improve robustness to certain factors, they induce spill-over effects to other factors. For example, color-jitter augmentation improves robustness to color and brightness, but surprisingly hurts robustness to pose. Together, these insights suggests that to advance the robustness of modern vision models, future research should focus on collecting additional diverse data and understanding data augmentation schemes. Along with these insights, we release a toolkit based on ImageNet-X to spur further study into the mistakes the image recognition systems make: https://facebookresearch.github.io/imagenetx/site/home .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0024c501-6412-4fb5-bbff-2292b5f39a92Cited by top-tier papers16
- The effectiveness of MAE pre-pretraining for billion-scale pretrainingMannat Singh, Quentin Duval, Kalyan Vasudev Alwala, Haoqi Fan et al.ICCV 2023 · 91 citations
- FACET: Fairness in Computer Vision Evaluation BenchmarkLaura Gustafson, Chloé Rolland, Nikhila Ravi, Quentin Duval et al.ICCV 2023 · 74 citations
- A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)Weijie Tu, Weijian Deng, Tom GedeonNeurIPS 2023 · 74 citations
- Understanding the detrimental class-level effects of data augmentationPolina Kirichenko, Mark Ibrahim, Randall Balestriero, Diane Bouchacourt et al.NeurIPS 2023 · 25 citations
- Identification of Systematic Errors of Image Classifiers on Rare SubgroupsJan Hendrik Metzen, Robin Hutmacher, N. Grace Hua, Valentyn Boreiko et al.ICCV 2023 · 23 citations
Builds on7
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini et al.NeurIPS 2020 · 731 citations
- Evaluating Machine Accuracy on ImageNetVaishaal Shankar, Rebecca Roelofs, Horia Mania, Alex Fang et al.ICML 2020 · 153 citations
- Salient ImageNet: How to discover spurious features in Deep Learning?Sahil Singla, Soheil FeiziICLR 2022 · 144 citations
- The Effects of Regularization and Data Augmentation are Class DependentRandall Balestriero, Léon Bottou, Yann LeCunNeurIPS 2022 · 124 citations
Related papers
- Unveiling AI's Blind Spots: An Oracle for In-Domain, Out-of-Domain, and Adversarial ErrorsShuangpeng Han, Mengmi ZhangICML 2025
- Contemplating Real-World Object ClassificationAli BorjiICLR 2021 · 2 citations
- ViewFool: Evaluating the Robustness of Visual Recognition to Adversarial ViewpointsYinpeng Dong, Shouwei Ruan, Hang Su, Caixin Kang et al.NeurIPS 2022 · 72 citations
- Improving robustness against common corruptions by covariate shift adaptationSteffen Schneider, Evgenia Rusak, Luisa Eck, Oliver Bringmann et al.NeurIPS 2020 · 688 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
