Towards Good Practices for Efficiently Annotating Large-Scale Image Classification Datasets
Yuan-Hong Liao, Amlan Kar, Sanja Fidler
摘要
Data is the engine of modern computer vision, which necessitates collecting large-scale datasets. This is expensive, and guaranteeing the quality of the labels is a major challenge. In this paper, we investigate efficient annotation strategies for collecting multi-class classification labels for a large collection of images. While methods that exploit learnt models for labeling exist, a surprisingly prevalent approach is to query humans for a fixed number of labels per datum and aggregate them, which is expensive. Building on prior work on online joint probabilistic modeling of human annotations and machine-generated beliefs, we propose modifications and best practices aimed at minimizing human labeling effort. Specifically, we make use of advances in self-supervised learning, view annotation as a semi-supervised learning problem, identify and mitigate pitfalls and ablate several key design choices to propose effective guidelines for labeling. Our analysis is done in a more realistic simulation that involves querying human labelers, which uncovers issues with evaluation using existing worker simulation methods. Simulated experiments on a 125k image subset of the ImageNet100 show that it can be annotated to 80% top-1 accuracy with 0.35 annotations per image on average, a 2.7x and 6.7x improvement over prior work and manual annotation, respectively. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu 等ICLR 2022 · 被引用 338 次
- Improving Contrastive Learning on Imbalanced Data via Open-World SamplingZiyu Jiang, Tianlong Chen, Ting Chen, Zhangyang WangNeurIPS 2021 · 被引用 52 次
- Learning to Defer with Limited Expert PredictionsPatrick Hemmer, Lukas Thede, Michael Vössing, Johannes Jakubik 等AAAI 2023 · 被引用 28 次
- Inside-Out: Measuring Generalization in Vision Transformers Through Inner WorkingsYunxiang Peng, Mengmeng Ma, Ziyu Yao, Xi PengCVPR 2026
它引用的顶会 Paper10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
相关 Paper
- Scaling up instance annotation via label propagationDim P. Papadopoulos, Ethan Weber, Antonio TorralbaICCV 2021 · 被引用 13 次
- From ImageNet to Image Classification: Contextualizing Progress on BenchmarksDimitris Tsipras, Shibani Santurkar, Logan Engstrom, Andrew Ilyas 等ICML 2020 · 被引用 146 次
- Similarity Search for Efficient Active Learning and Search of Rare ConceptsCody Coleman, Edward Chou, Julian Katz-Samuels, Sean Culatana 等AAAI 2022 · 被引用 47 次
- MCAL: Minimum Cost Human-Machine Active LabelingHang Qiu, Krishna Chintalapudi, Ramesh GovindanICLR 2023
- Active Learning for Semantic Segmentation with Multi-class Label QuerySehyun Hwang, Sohyun Lee, Hoyoung Kim, Minhyeon Oh 等NeurIPS 2023 · 被引用 22 次
