Scaling Up Active Testing to Large Language Models
Gabrielle Berrada, Jannik Kossen, Freddie Bickford Smith, Muhammed Razzak, Yarin Gal, Thomas Rainforth
Abstract
Active testing enables label-efficient evaluation of predictive models through careful data acquisition, but it can pose a significant computational cost. We identify cost-saving measures that enable active testing to be scaled up to large language models (LLMs). In particular we show that the surrogate model used to guide data acquisition can be constructed cheaply using in-context learning, does not require updating within an active-testing loop, and can be smaller than the target model. We even find we can make good data-acquisition decisions without making predictions with the target model. As a result we are able to achieve much more accurate evaluations of LLM performance relative to using randomly acquired data. We additionally introduce a bootstrap estimator of evaluation error, which we show to be a useful indicator of how well active testing is working within a single run.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9f0f003-6791-4a5d-a22a-e5d1e3ce3ecaCited by top-tier papers1
Ask how each one uses itBuilds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- On Statistical Bias In Active Learning: How and When to Fix ItSebastian Farquhar, Yarin Gal, Tom RainforthICLR 2021 · 96 citations
- Active Testing: Sample-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Tom RainforthICML 2021 · 81 citations
- In-Context Learning Learns Label Relationships but Is Not Conventional LearningJannik Kossen, Yarin Gal, Tom RainforthICLR 2024 · 61 citations
Related papers
- Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Thomas RainforthNeurIPS 2022 · 36 citations
- Probing the Decision Boundaries of In-context Learning in Large Language ModelsSiyan Zhao, Tung Nguyen, Aditya GroverNeurIPS 2024 · 26 citations
- Active Learning with LLMs for Partially Observed and Cost-Aware ScenariosNicolás Astorga, Tennison Liu, Nabeel Seedat, Mihaela van der SchaarNeurIPS 2024 · 11 citations
- Active Evaluation Acquisition for Efficient LLM BenchmarkingYang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba et al.ICML 2025
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 128 citations
