MCAL: Minimum Cost Human-Machine Active Labeling
Hang Qiu, Krishna Chintalapudi, Ramesh Govindan
摘要
Today, ground-truth generation uses data sets annotated by cloud-based annotation services. These services rely on human annotation, which can be prohibitively expensive. In this paper, we consider the problem of hybrid human-machine labeling, which trains a classifier to accurately auto-label part of the data set. However, training the classifier can be expensive too. We propose an iterative approach that minimizes total overall cost by, at each step, jointly determining which samples to label using humans and which to label using the trained classifier. We validate our approach on well known public data sets such as Fashion-MNIST, CIFAR-10, CIFAR-100, and ImageNet. In some cases, our approach has 6× lower overall cost relative to human labeling the entire data set, and is always cheaper than the cheapest competing strategy. INTRODUCTION Ground-truth is crucial for training and testing ML models. Generating accurate ground-truth was cumbersome until the recent emergence of cloud-based human annotation services (SageMaker (2021); Google (2021); Figure-Eight (2021)). Users of these services submit data sets and receive, in return, annotations on each data item in the data set. Because these services typically employ humans to generate ground-truth, annotation costs can be prohibitively high especially for large data sets. Hybrid Human-machine Annotations. In this paper, we explore using a hybrid human-machine approach to reduce annotation costs (in $) where humans only annotate a subset of the data items; a machine learning model trained on this annotated data annotates the rest. The accuracy of a model trained on a subset of the data set will typically be lower than that of human annotators. However, a user of an annotation service might choose to avail of this trade-off if (a) targeting a slightly lower annotation quality can significantly reduce costs, or (b) the cost of training a model to a higher accuracy is itself prohibitive. Consequently, this paper focuses on the design of a hybrid human-machine annotation scheme that minimizes the overall cost of annotating the entire data set (including the cost of training the model) while ensuring that the overall annotation accuracy, relative to human annotations, is higher than a pre-specified target (e.g., 95%). Challenges. In this paper, we consider a specific annotation task, multi-class labeling. We assume that the user of an annotation service provides a set X of data to be labeled and a classifier D to use for machine labeling. Then, the goal is to find a subset B ⊂ X human-labeled samples to train D, and use the classifier to label the rest, minimizing total cost while ensuring the target accuracy. A straw man approach might seek to predict human-labeled subset B in a single shot. This is hard to do because it depends on several factors: (a) the classifier architecture and how much accuracy it can achieve, (b) how "hard" the dataset is, (c) the cost of training and labeling, and (d) the target accuracy. Complex models may provide a high accuracy, their training costs may be too high and potentially offset the gains obtained through machine-generated annotations. Some data-points in a dataset are more informative as compared to the rest from a model training perspective. Identifying the "right" data subset for human-vs. machine-labeling can minimize the total labeling cost. Approach. In this paper we propose a novel technique, MCAL 1 (Minimum Cost Active Labeling), that addresses these challenges and is able to minimize annotation cost across diverse data sets. At its
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Provably Label-Efficient Conformal PredictionAndrew Ilyas, Joonhyuk Ko, Jingwu Tang, Steven Wu 等ICML 2026 · 被引用 14 次
- Probably Approximately Correct LabelsEmmanuel J Candes, Andrew Ilyas, Tijana ZrnicICML 2026 · 被引用 7 次
- Pearls from Pebbles: Improved Confidence Functions for Auto-labelingHarit Vishwakarma, Yi Chen, Sui Jiet Tay, Satya Sai Srinath Namburi 等NeurIPS 2024 · 被引用 7 次
- SciLitLLM: How to Adapt LLMs for Scientific Literature UnderstandingSihang Li, Jin Huang, Jiaxi Zhuang, Yaorui Shi 等ICLR 2025 · 被引用 1 次
它引用的顶会 Paper3
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford 等ICLR 2020 · 被引用 974 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
- Adaptive Region-Based Active LearningCorinna Cortes, Giulia DeSalvo, Claudio Gentile, Mehryar Mohri 等ICML 2020 · 被引用 24 次
相关 Paper
- Towards Good Practices for Efficiently Annotating Large-Scale Image Classification DatasetsYuan-Hong Liao, Amlan Kar, Sanja FidlerCVPR 2021
- Learning a Cost-Effective Annotation Policy for Question AnsweringBernhard Kratzwald, Stefan Feuerriegel, Huan SunEMNLP 2020 · 被引用 9 次
- Double-Cross Attacks: Subverting Active Learning SystemsJose Rodrigo Sanchez Vicarte, Gang Wang, Christopher W. FletcherUSENIX Security 2021 · 被引用 8 次
- COCA: Cost-Effective Collaborative Annotation System by Combining Experts and AmateursJiayu Lei, Zheng Zhang, Lan Zhang, Xiang-Yang LiICDE 2022 · 被引用 6 次
- Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active LearningWenjun Qiu, David Lie, Lisa M. AustinUSENIX Security 2023
