Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget
Florian E. Dorner, Moritz Hardt
Abstract
We study how to best spend a budget of noisy labels to compare the accuracy of two binary classifiers. It's common practice to collect and aggregate multiple noisy labels for a given data point into a less noisy label via a majority vote. We prove a theorem that runs counter to conventional wisdom. If the goal is to identify the better of two classifiers, we show it's best to spend the budget on collecting a single label for more samples. Our result follows from a non-trivial application of Cramér's theorem, a staple in the theory of large deviations. We discuss the implications of our work for the design of machine learning benchmarks, where they overturn some time-honored recommendations. In addition, our results provide sample size bounds superior to what follows from Hoeffding's bound.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 13ef9ed6-4da9-477a-be5f-1c1618f06e85Cited by top-tier papers2
- Lawma: The Power of Specialization for Legal AnnotationRicardo Dominguez-Olmedo, Vedant Nanda, Rediet Abebe, Stefan Bechtold et al.ICLR 2025 · 1 citation
- Limits to scalable evaluation at the frontier: LLM as judge won't beat twice the dataFlorian E. Dorner, Vivian Yvonne Nastl, Moritz HardtICLR 2025
Builds on5
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Jury Learning: Integrating Dissenting Voices into Machine Learning ModelsMitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel et al.CHI 2022 · 134 citations
- To Aggregate or Not? Learning with Separate Noisy LabelsJiaheng Wei, Zhaowei Zhu, Tianyi Luo, Ehsan Amid et al.KDD 2023 · 21 citations
- FairPrism: Evaluating Fairness-Related Harms in Text GenerationEve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett et al.ACL 2023 · 9 citations
- Human-Guided Fair Classification for Natural Language ProcessingFlorian E. Dorner, Momchil Peychev, Nikola Konstantinov, Naman Goel et al.ICLR 2023
Related papers
- Cost-Accuracy Aware Adaptive Labeling for Active LearningRuijiang Gao, Maytal Saar-TsechanskyAAAI 2020 · 23 citations
- The Price of Fairness in Active Learning: Fundamental Limits and Optimal Label AcquisitionChang Lu, Yizheng ZhaoKDD 2026
- The Many Faces of Optimal Weak-to-Strong LearningMikael Møller Høgsgaard, Kasper Green Larsen, Markus Engelund MathiasenNeurIPS 2024 · 4 citations
- PAC-Bayesian Bounds on Rate-Efficient ClassifiersAlhabib Abbas, Yiannis AndreopoulosICML 2022 · 1 citation
- Second Order PAC-Bayesian Bounds for the Weighted Majority VoteAndrés R. Masegosa, Stephan Sloth Lorenzen, Christian Igel, Yevgeny SeldinNeurIPS 2020 · 48 citations
