Automated Classification of Model Errors on ImageNet
Momchil Peychev, Mark Niklas Müller, Marc Fischer, Martin T. Vechev
摘要
While the ImageNet dataset has been driving computer vision research over the past decade, significant label noise and ambiguity have made top-1 accuracy an insufficient measure of further progress. To address this, new label-sets and evaluation protocols have been proposed for ImageNet showing that state-of-the-art models already achieve over 95% accuracy and shifting the focus on investigating why the remaining errors persist. Recent work in this direction employed a panel of experts to manually categorize all remaining classification errors for two selected models. However, this process is time-consuming, prone to inconsistencies, and requires trained experts, making it unsuitable for regular model evaluation thus limiting its utility. To overcome these limitations, we propose the first automated error classification framework, a valuable tool to study how modeling choices affect error distributions. We use our framework to comprehensively evaluate the error distribution of over 900 models. Perhaps surprisingly, we find that across model architectures, scales, and pre-training corpora, top-1 accuracy is a strong predictor for the portion of all error types. In particular, we observe that the portion of severe errors drops significantly with top-1 accuracy indicating that, while it underreports a model's true performance, it remains a valuable performance metric. We release all our code at https://github.com/eth-sri/automated-error-analysis .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Efficient Lifelong Model Evaluation in an Era of Rapid ProgressAmeya Prabhu, Vishaal Udandarao, Philip Torr, Matthias Bethge 等NeurIPS 2024 · 被引用 11 次
- On Large Multimodal Models as Open-World Image ClassifiersAlessandro Conti, Massimiliano Mancini, Enrico Fini, Yiming Wang 等ICCV 2025 · 被引用 3 次
- Fine-Grained Erasure in Text-to-Image Diffusion-based Foundation ModelsKartik Thakral, Tamar Glaser, Tal Hassner, Mayank Vatsa 等CVPR 2025
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Do Wide and Deep Networks Learn the Same Things? Uncovering How Neural Network Representations Vary with Width and DepthThao Nguyen, Maithra Raghu, Simon KornblithICLR 2021 · 被引用 323 次
相关 Paper
- When does dough become a bagel? Analyzing the remaining mistakes on ImageNetVijay Vasudevan, Benjamin Caine, Raphael Gontijo Lopes, Sara Fridovich-Keil 等NeurIPS 2022 · 被引用 79 次
- Re-Labeling ImageNet: From Single to Multi-Labels, From Global to Localized LabelsSangdoo Yun, Seong Joon Oh, Byeongho Heo, Dongyoon Han 等CVPR 2021
- Evaluating Machine Accuracy on ImageNetVaishaal Shankar, Rebecca Roelofs, Horia Mania, Alex Fang 等ICML 2020 · 被引用 153 次
- From ImageNet to Image Classification: Contextualizing Progress on BenchmarksDimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas 等ICML 2020 · 被引用 146 次
- Salient ImageNet: How to discover spurious features in Deep Learning?Sahil Singla, Soheil FeiziICLR 2022 · 被引用 144 次
