The ALCHEmist: Automated Labeling 500x CHEaper than LLM Data Annotators
Tzu-Heng Huang, Catherine Cao, Vaishnavi Bhargava, Frederic Sala
Abstract
Large pretrained models can be used as annotators, helping replace or augment crowdworkers and enabling distilling generalist models into smaller specialist models. Unfortunately, this comes at a cost: employing top-of-the-line models often requires paying thousands of dollars for API calls, while the resulting datasets are static and challenging to audit. To address these challenges, we propose a simple alternative: rather than directly querying labels from pretrained models, we task models to generate programs that can produce labels. These programs can be stored and applied locally, re-used and extended, and cost orders of magnitude less. Our system, Alchemist, obtains comparable to or better performance than large language model-based annotation in a range of tasks for a fraction of the cost: on average, improvements amount to a 12.9% enhancement while the total labeling costs across all datasets are reduced by a factor of approximately 500x.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 731c0219-06f8-4d82-a627-aada3988183fCited by top-tier papers3
- CARE: Confounder-Aware Aggregation for Reliable LLM EvaluationJitian Zhao, Changho Shin, Tzu-Heng Huang, Satya Sai Srinath Namburi GNVV et al.ICML 2026 · 7 citations
- Ensembling LLM-Induced Decision Trees for Explainable and Robust Error DetectionMengqi Wang, Jianwei Wang, Qing Liu, Xiwei Xu et al.KDD 2026 · 4 citations
- Evaluating Sample Utility for Efficient Data Selection by Mimicking Model WeightsTzu-Heng Huang, Manjot Bilkhu, John Cooper, Frederic Sala et al.ICML 2026 · 3 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng et al.EMNLP 2024 · 479 citations
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper et al.ICML 2020 · 130 citations
Related papers
- AlchemistCoder: Harmonizing and Eliciting Code Capability by Hindsight Tuning on Multi-source DataZifan Song, Yudong Wang, Wenwei Zhang, Kuikun Liu et al.NeurIPS 2024 · 8 citations
- Prompting Is Programming: A Query Language for Large Language ModelsLuca Beurer-Kellner, Marc Fischer, Martin T. VechevPLDI 2023 · 114 citations
- Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data AnnotationMingxuan Xia, Haobo Wang, Yixuan Li, Zewei Yu et al.ACL 2025 · 4 citations
- Generating Data for Symbolic Language with Large Language ModelsJiacheng Ye, Chengzu Li, Lingpeng Kong, Tao YuEMNLP 2023 · 8 citations
- Interactive Multi-fidelity Learning for Cost-effective Adaptation of Language Model with Sparse Human SupervisionJiaxin Zhang, Zhuohang Li, Kamalika Das, Kumar SricharanNeurIPS 2023 · 6 citations
