Promises and Pitfalls of Threshold-based Auto-labeling
Harit Vishwakarma, Heguang Lin, Frederic Sala, Ramya Korlakai Vinayak
摘要
Creating large-scale high-quality labeled datasets is a major bottleneck in supervised machine learning workflows. Threshold-based auto-labeling (TBAL), where validation data obtained from humans is used to find a confidence threshold above which the data is machine-labeled, reduces reliance on manual annotation. TBAL is emerging as a widely-used solution in practice. Given the long shelf-life and diverse usage of the resulting datasets, understanding when the data obtained by such auto-labeling systems can be relied on is crucial. This is the first work to analyze TBAL systems and derive sample complexity bounds on the amount of human-labeled validation data required for guaranteeing the quality of machine-labeled data. Our results provide two crucial insights. First, reasonable chunks of unlabeled data can be automatically and accurately labeled by seemingly bad models. Second, a hidden downside of TBAL systems is potentially prohibitive validation data usage. Together, these insights describe the promise and pitfalls of using such systems. We validate our theoretical guarantees with extensive experiments on synthetic and real datasets. 1 2
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Active Statistical InferenceTijana Zrnic, Emmanuel J. CandèsICML 2024 · 被引用 34 次
- Pearls from Pebbles: Improved Confidence Functions for Auto-labelingHarit Vishwakarma, Yi Chen, Sui Jiet Tay, Satya Sai Srinath Namburi 等NeurIPS 2024 · 被引用 7 次
- Query Design for Crowdsourced Clustering: Effect of Cognitive Overload and Contextual BiasYi Chen, Ramya Korlakai VinayakWWW 2025 · 被引用 3 次
- Online Adaptive Anomaly Thresholding with Confidence SequencesSophia Huiwen Sun, Abishek Sankararaman, Balakrishnan NarayanaswamyICML 2024 · 被引用 2 次
- ULAREF: A Unified Label Refinement Framework for Learning with Inaccurate SupervisionCongyu Qiao, Ning Xu, Yihao Hu, Xin GengICML 2024 · 被引用 1 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis 等NeurIPS 2021 · 被引用 633 次
- Batch Active Learning at ScaleGui Citovsky, Giulia DeSalvo, Claudio Gentile, Lazaros Karydas 等NeurIPS 2021 · 被引用 220 次
- Improving model calibration with accuracy versus uncertainty optimizationRanganath Krishnan, Omesh TickooNeurIPS 2020 · 被引用 217 次
- Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of OverconfidenceDeng-Bao Wang, Lei Feng, Min-Ling ZhangNeurIPS 2021 · 被引用 177 次
相关 Paper
- AutoEval Done Right: Using Synthetic Data for Model EvaluationPierre Boyeau, Anastasios Nikolas Angelopoulos, Tianle Li, Nir Yosef 等ICML 2025
- LANCET: Labeling Complex Data at ScaleHuayi Zhang, Lei Cao, Samuel Madden, Elke A. RundensteinerVLDB 2021 · 被引用 12 次
- MCAL: Minimum Cost Human-Machine Active LabelingHang Qiu, Krishna Chintalapudi, Ramesh GovindanICLR 2023
- Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question AnsweringJian Lan, Zhicheng Liu, Udo Schlegel, Raoyuan Zhao 等ICLR 2026 · 被引用 2 次
- On Efficient and Statistical Quality Estimation for Data AnnotationJan-Christoph Klie, Juan Haladjian, Marc Kirchner, Rahul NairACL 2024 · 被引用 1 次
