Promises and Pitfalls of Threshold-based Auto-labeling
Harit Vishwakarma, Heguang Lin, Frederic Sala, Ramya Korlakai Vinayak
Abstract
Creating large-scale high-quality labeled datasets is a major bottleneck in supervised machine learning workflows. Threshold-based auto-labeling (TBAL), where validation data obtained from humans is used to find a confidence threshold above which the data is machine-labeled, reduces reliance on manual annotation. TBAL is emerging as a widely-used solution in practice. Given the long shelf-life and diverse usage of the resulting datasets, understanding when the data obtained by such auto-labeling systems can be relied on is crucial. This is the first work to analyze TBAL systems and derive sample complexity bounds on the amount of human-labeled validation data required for guaranteeing the quality of machine-labeled data. Our results provide two crucial insights. First, reasonable chunks of unlabeled data can be automatically and accurately labeled by seemingly bad models. Second, a hidden downside of TBAL systems is potentially prohibitive validation data usage. Together, these insights describe the promise and pitfalls of using such systems. We validate our theoretical guarantees with extensive experiments on synthetic and real datasets. 1 2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1dda484-ba8e-41e5-920b-e774e9981060Cited by top-tier papers7
- Active Statistical InferenceTijana Zrnic, Emmanuel J. CandèsICML 2024 · 34 citations
- Pearls from Pebbles: Improved Confidence Functions for Auto-labelingHarit Vishwakarma, Yi Chen, Sui Jiet Tay, Satya Sai Srinath Namburi et al.NeurIPS 2024 · 7 citations
- Query Design for Crowdsourced Clustering: Effect of Cognitive Overload and Contextual BiasYi Chen, Ramya Korlakai VinayakWWW 2025 · 3 citations
- Online Adaptive Anomaly Thresholding with Confidence SequencesSophia Huiwen Sun, Abishek Sankararaman, Balakrishnan NarayanaswamyICML 2024 · 2 citations
- ULAREF: A Unified Label Refinement Framework for Learning with Inaccurate SupervisionCongyu Qiao, Ning Xu, Yihao Hu, Xin GengICML 2024 · 1 citation
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis et al.NeurIPS 2021 · 633 citations
- Batch Active Learning at ScaleGui Citovsky, Giulia DeSalvo, Claudio Gentile, Lazaros Karydas et al.NeurIPS 2021 · 220 citations
- Improving model calibration with accuracy versus uncertainty optimizationRanganath Krishnan, Omesh TickooNeurIPS 2020 · 217 citations
- Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of OverconfidenceDeng-Bao Wang, Lei Feng, Min-Ling ZhangNeurIPS 2021 · 177 citations
Related papers
- AutoEval Done Right: Using Synthetic Data for Model EvaluationPierre Boyeau, Anastasios Nikolas Angelopoulos, Tianle Li, Nir Yosef et al.ICML 2025
- LANCET: Labeling Complex Data at ScaleHuayi Zhang, Lei Cao, Samuel Madden, Elke A. RundensteinerVLDB 2021 · 12 citations
- MCAL: Minimum Cost Human-Machine Active LabelingHang Qiu, Krishna Chintalapudi, Ramesh GovindanICLR 2023
- Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question AnsweringJian Lan, Zhicheng Liu, Udo Schlegel, Raoyuan Zhao et al.ICLR 2026 · 2 citations
- On Efficient and Statistical Quality Estimation for Data AnnotationJan-Christoph Klie, Juan Haladjian, Marc Kirchner, Rahul NairACL 2024 · 1 citation
