From ImageNet to Image Classification: Contextualizing Progress on Benchmarks
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas, Aleksander Madry
Abstract
Building rich machine learning datasets in a scalable manner often necessitates a crowd-sourced data collection pipeline. In this work, we use human studies to investigate the consequences of employing such a pipeline, focusing on the popular ImageNet dataset. We study how specific design choices in the ImageNet creation process impact the fidelity of the resulting dataset-including the introduction of biases that state-of-the-art models exploit. Our analysis pinpoints how a noisy data collection pipeline can lead to a systematic misalignment between the resulting benchmark and the real-world task it serves as a proxy for. Finally, our findings emphasize the need to augment our current model training and evaluation toolkit to take such misalignments into account. 1 * Equal contribution. 1 To facilitate further research, we release our refined ImageNet annotations at https://github.com/MadryLab/ ImageNetMultiLabel .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7c0e6a17-ef0b-45f7-b078-96ac7ea61967Cited by top-tier papers30
- BREEDS: Benchmarks for Subpopulation ShiftShibani Santurkar, Dimitris Tsipras, Aleksander MadryICLR 2021 · 193 citations
- Salient ImageNet: How to discover spurious features in Deep Learning?Sahil Singla, Soheil FeiziICLR 2022 · 144 citations
- Editing a classifier by rewriting its prediction rulesShibani Santurkar, Dimitris Tsipras, Mahalaxmi Elango, David Bau et al.NeurIPS 2021 · 105 citations
- Leveraging Sparse Linear Layers for Debuggable Deep NetworksEric Wong, Shibani Santurkar, Aleksander MadryICML 2021 · 101 citations
- When does dough become a bagel? Analyzing the remaining mistakes on ImageNetVijay Vasudevan, Benjamin Caine, Raphael Gontijo Lopes, Sara Fridovich-Keil et al.NeurIPS 2022 · 79 citations
Builds on2
Related papers
- Towards Good Practices for Efficiently Annotating Large-Scale Image Classification DatasetsYuan-Hong Liao, Amlan Kar, Sanja FidlerCVPR 2021
- Impact of Annotator Demographics on Sentiment Dataset LabelingYi Ding, Jacob You, Tonja-Katrin Machulla, Jennifer Jacobs et al.CSCW 2022 · 20 citations
- Re-Labeling ImageNet: From Single to Multi-Labels, From Global to Localized LabelsSangdoo Yun, Seong Joon Oh, Byeongho Heo, Dongyoon Han et al.CVPR 2021
- Evaluating Machine Accuracy on ImageNetVaishaal Shankar, Rebecca Roelofs, Horia Mania, Alex Fang et al.ICML 2020 · 153 citations
- Detecting Label Errors by Using Pre-Trained Language ModelsDerek Chong, Jenny Hong, Christopher D. ManningEMNLP 2022 · 8 citations
