Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets
Yuan-Hong Liao, Amlan Kar, Sanja Fidler
Abstract
Data is the engine of modern computer vision, which necessitates collecting large-scale datasets. This is expensive, and guaranteeing the quality of the labels is a major challenge. In this paper, we investigate efficient annotation strategies for collecting multi-class classification labels for a large collection of images. While methods that exploit learnt models for labeling exist, a surprisingly prevalent approach is to query humans for a fixed number of labels per datum and aggregate them, which is expensive. Building on prior work on online joint probabilistic modeling of human annotations and machine-generated beliefs, we propose modifications and best practices aimed at minimizing human labeling effort. Specifically, we make use of advances in self-supervised learning, view annotation as a semi-supervised learning problem, identify and mitigate pitfalls and ablate several key design choices to propose effective guidelines for labeling. Our analysis is done in a more realistic simulation that involves querying human labelers, which uncovers issues with evaluation using existing worker simulation methods. Simulated experiments on a 125k image subset of the ImageNet100 show that it can be annotated to 80% top-1 accuracy with 0.35 annotations per image on average, a 2.7x and 6.7x improvement over prior work and manual annotation, respectively. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 078e3bc8-23cc-4683-a399-4e516e20aa5dCited by top-tier papers4
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu et al.ICLR 2022 · 338 citations
- Improving Contrastive Learning on Imbalanced Data via Open-World SamplingZiyu Jiang, Tianlong Chen, Ting Chen, Zhangyang WangNeurIPS 2021 · 52 citations
- Learning to Defer with Limited Expert PredictionsPatrick Hemmer, Lukas Thede, Michael Vössing, Johannes Jakubik et al.AAAI 2023 · 28 citations
- Inside-Out: Measuring Generalization in Vision Transformers Through Inner WorkingsYunxiang Peng, Mengmeng Ma, Ziyu Yao, Xi PengCVPR 2026
Builds on10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
Related papers
- Scaling up instance annotation via label propagationDim P. Papadopoulos, Ethan Weber, Antonio TorralbaICCV 2021 · 13 citations
- From ImageNet to Image Classification: Contextualizing Progress on BenchmarksDimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas et al.ICML 2020 · 146 citations
- Similarity Search for Efficient Active Learning and Search of Rare ConceptsCody Coleman, Edward Chou, Julian Katz-Samuels, Sean Culatana et al.AAAI 2022 · 47 citations
- MCAL: Minimum Cost Human-Machine Active LabelingHang Qiu, Krishna Chintalapudi, Ramesh GovindanICLR 2023
- Active Learning for Semantic Segmentation with Multi-class Label QuerySehyun Hwang, Sohyun Lee, Hoyoung Kim, Minhyeon Oh et al.NeurIPS 2023 · 22 citations
