Probably Approximately Correct Labels
Emmanuel J Candes, Andrew Ilyas, Tijana Zrnic
摘要
Obtaining high-quality labeled datasets is often costly, requiring either human annotation or expensive experiments. In theory, powerful pre-trained AI models provide an opportunity to automatically label datasets and save costs. Unfortunately, these models come with no guarantees on their accuracy, making wholesale replacement of manual labeling impractical. In this work, we propose a method for leveraging pre-trained AI models to curate cost-effective and high-quality datasets. In particular, our approach results in probably approximately correct labels: with high probability, the overall labeling error is small. Our method is nonasymptotically valid under minimal assumptions on the dataset or the AI model being studied, and thus enables rigorous yet efficient dataset curation using modern AI models. We demonstrate the benefits of the methodology through text annotation with large language models, image labeling with pre-trained vision models, and protein folding analysis with AlphaFold.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Provably Label-Efficient Conformal PredictionAndrew Ilyas, Joonhyuk Ko, Jingwu Tang, Steven Wu 等ICML 2026 · 被引用 14 次
- Anytime Safe PAC Efficient ReasoningChengyao Yu, Hao Zeng, Youxin Zhu, Jianguo Huang 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper8
- Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language ModelsNaoki Egami, Musashi Hinck, Brandon M. Stewart, Hanying WeiNeurIPS 2023 · 被引用 74 次
- Active Statistical InferenceTijana Zrnic, Emmanuel J. CandèsICML 2024 · 被引用 34 次
- CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data AnnotationMinzhi Li, Taiwei Shi, Caleb Ziems, Min-Yen Kan 等EMNLP 2023 · 被引用 33 次
- We Can Detect Your Bias: Predicting the Political Ideology of News ArticlesRamy Baly, Giovanni Da San Martino, James R. Glass, Preslav NakovEMNLP 2020 · 被引用 6 次
- Misinfo Reaction Frames: Reasoning about Readers' Reactions to News HeadlinesSaadia Gabriel, Skyler Hallinan, Maarten Sap, Pemi Nguyen 等ACL 2022
相关 Paper
- The ALCHEmist: Automated Labeling 500x CHEaper than LLM Data AnnotatorsTzu-Heng Huang, Catherine Cao, Vaishnavi Bhargava, Frederic SalaNeurIPS 2024 · 被引用 15 次
- SHED: Shapley-Based Automated Dataset Refinement for Instruction Fine-TuningYexiao He, Ziyao Wang, Zheyu Shen, Guoheng Sun 等NeurIPS 2024 · 被引用 24 次
- One protein is all you needAnton Bushuiev, Roman Bushuiev, Olga Pimenova, Nikola Zadorozhny 等ICLR 2026 · 被引用 1 次
- Escaping Collapse: The Strength of Weak Data for Large Language Model TrainingKareem Amin, Sara Babakniya, Alex Bie, Weiwei Kong 等NeurIPS 2025 · 被引用 17 次
- CiT: Curation in Training for Effective Vision-Language DataHu Xu, Saining Xie, Po-Yao Huang, Licheng Yu 等ICCV 2023 · 被引用 31 次
