Pearls from Pebbles: Improved Confidence Functions for Auto-labeling
Harit Vishwakarma, Yi Chen, Sui Jiet Tay, Satya Sai Srinath Namburi, Frederic Sala, Ramya Korlakai Vinayak
摘要
Auto-labeling is an important family of techniques that produce labeled training sets with minimum manual labeling. A prominent variant, threshold-based auto-labeling (TBAL), works by finding a threshold on a model's confidence scores above which it can accurately label unlabeled data points. However, many models are known to produce overconfident scores, leading to poor TBAL performance. While a natural idea is to apply off-the-shelf calibration methods to alleviate the overconfidence issue, such methods still fall short. Rather than experimenting with ad-hoc choices of confidence functions, we propose a framework for studying the optimal TBAL confidence function. We develop a tractable version of the framework to obtain Colander (Confidence functions for Efficient and Reliable Auto-labeling), a new post-hoc method specifically designed to maximize performance in TBAL systems. We perform an extensive empirical evaluation of our method Colander and compare it against methods designed for calibration. Colander achieves up to 60% improvements on coverage over the baselines while maintaining auto-labeling error below and using the same amount of labeled data as the baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Calibrating Deep Neural Networks using Focal LossJishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz 等NeurIPS 2020 · 被引用 674 次
- Confidence-Aware Learning for Deep Neural NetworksJooyoung Moon, Jihyo Kim, Younghak Shin, Sangheum HwangICML 2020 · 被引用 184 次
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper 等ICML 2020 · 被引用 130 次
相关 Paper
- Promises and Pitfalls of Threshold-based Auto-labelingHarit Vishwakarma, Heguang Lin, Frederic Sala, Ramya Korlakai VinayakNeurIPS 2023 · 被引用 16 次
- LANCET: Labeling Complex Data at ScaleHuayi Zhang, Lei Cao, Samuel Madden, Elke A. RundensteinerVLDB 2021 · 被引用 12 次
- LaSCal: Label-Shift Calibration without target labelsTeodora Popordanoska, Gorjan Radevski, Tinne Tuytelaars, Matthew B. BlaschkoNeurIPS 2024 · 被引用 12 次
- Top-label calibration and multiclass-to-binary reductionsChirag Gupta, Aaditya RamdasICLR 2022 · 被引用 51 次
- Rethinking Calibration of Deep Neural Networks: Do Not Be Afraid of OverconfidenceDeng-Bao Wang, Lei Feng, Min-Ling ZhangNeurIPS 2021 · 被引用 177 次
