Probably Approximately Correct Labels
Emmanuel J Candes, Andrew Ilyas, Tijana Zrnic
Abstract
Obtaining high-quality labeled datasets is often costly, requiring either human annotation or expensive experiments. In theory, powerful pre-trained AI models provide an opportunity to automatically label datasets and save costs. Unfortunately, these models come with no guarantees on their accuracy, making wholesale replacement of manual labeling impractical. In this work, we propose a method for leveraging pre-trained AI models to curate cost-effective and high-quality datasets. In particular, our approach results in probably approximately correct labels: with high probability, the overall labeling error is small. Our method is nonasymptotically valid under minimal assumptions on the dataset or the AI model being studied, and thus enables rigorous yet efficient dataset curation using modern AI models. We demonstrate the benefits of the methodology through text annotation with large language models, image labeling with pre-trained vision models, and protein folding analysis with AlphaFold.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43c34361-aa19-4366-84db-9f3ae2a3ad6bCited by top-tier papers2
- Provably Label-Efficient Conformal PredictionAndrew Ilyas, Joonhyuk Ko, Jingwu Tang, Steven Wu et al.ICML 2026 · 14 citations
- Anytime Safe PAC Efficient ReasoningChengyao Yu, Hao Zeng, Youxin Zhu, Jianguo Huang et al.ICML 2026 · 2 citations
Builds on8
- Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language ModelsNaoki Egami, Musashi Hinck, Brandon M. Stewart, Hanying WeiNeurIPS 2023 · 74 citations
- Active Statistical InferenceTijana Zrnic, Emmanuel J. CandèsICML 2024 · 34 citations
- CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data AnnotationMinzhi Li, Taiwei Shi, Caleb Ziems, Min-Yen Kan et al.EMNLP 2023 · 33 citations
- We Can Detect Your Bias: Predicting the Political Ideology of News ArticlesRamy Baly, Giovanni Da San Martino, James R. Glass, Preslav NakovEMNLP 2020 · 6 citations
- Misinfo Reaction Frames: Reasoning about Readers' Reactions to News HeadlinesSaadia Gabriel, Skyler Hallinan, Maarten Sap, Pemi Nguyen et al.ACL 2022
Related papers
- The ALCHEmist: Automated Labeling 500x CHEaper than LLM Data AnnotatorsTzu-Heng Huang, Catherine Cao, Vaishnavi Bhargava, Frederic SalaNeurIPS 2024 · 15 citations
- SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-TuningYexiao He, Ziyao Wang, Zheyu Shen, Guoheng Sun et al.NeurIPS 2024 · 24 citations
- One protein is all you needAnton Bushuiev, Roman Bushuiev, Olga Pimenova, Nikola Zadorozhny et al.ICLR 2026 · 1 citation
- Escaping Collapse: The Strength of Weak Data for Large Language Model TrainingKareem Amin, Sara Babakniya, Alex Bie, Weiwei Kong et al.NeurIPS 2025 · 17 citations
- CiT: Curation in Training for Effective Vision-Language DataHu Xu, Saining Xie, Po-Yao Huang, Licheng Yu et al.ICCV 2023 · 31 citations
