Active Testing: Sample-Efficient Model Evaluation
Jannik Kossen, Sebastian Farquhar, Yarin Gal, Tom Rainforth
摘要
We introduce a new framework for sample-efficient model evaluation that we call active testing. While approaches like active learning reduce the number of labels needed for model training, existing literature largely ignores the cost of labeling test data, typically unrealistically assuming large test sets for model evaluation. This creates a disconnect to real applications, where test labels are important and just as expensive, e.g. for optimizing hyperparameters. Active testing addresses this by carefully selecting the test points to label, ensuring model evaluation is sample-efficient. To this end, we derive theoretically-grounded and intuitive acquisition strategies that are specifically tailored to the goals of active testing, noting these are distinct to those of active learning. As actively selecting labels introduces a bias; we further show how to remove this bias while reducing the variance of the estimator at the same time. Active testing is easy to implement and can be applied to any supervised machine learning method. We demonstrate its effectiveness on models including WideResNets and Gaussian processes on datasets including Fashion-MNIST and CIFAR-100.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun 等ICML 2024 · 被引用 212 次
- Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational DataAndrew Jesson, Panagiotis Tigas, Joost van Amersfoort, Andreas Kirsch 等NeurIPS 2021 · 被引用 42 次
- Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Thomas RainforthNeurIPS 2022 · 被引用 36 次
- Navigating the Pitfalls of Active Learning Evaluation: A Systematic Framework for Meaningful Performance AssessmentCarsten T. Lüth, Till J. Bungert, Lukas Klein, Paul F. JaegerNeurIPS 2023 · 被引用 32 次
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 被引用 26 次
它引用的顶会 Paper3
- Deep Adaptive Design: Amortizing Sequential Bayesian Experimental DesignAdam Foster, Desi R. Ivanova, Ilyas Malik, Tom RainforthICML 2021 · 被引用 119 次
- On Statistical Bias In Active Learning: How and When to Fix ItSebastian Farquhar, Yarin Gal, Tom RainforthICLR 2021 · 被引用 96 次
- Active Bayesian Assessment of Black-Box ClassifiersDisi Ji, Robert L. Logan IV, Padhraic Smyth, Mark SteyversAAAI 2021 · 被引用 3 次
相关 Paper
- Actively Testing Your Model While It Learns: Realizing Label-Efficient Learning in PracticeDayou Yu, Weishi Shi, Qi YuNeurIPS 2023 · 被引用 5 次
- Scaling Up Active Testing to Large Language ModelsGabrielle Berrada, Jannik Kossen, Freddie Bickford Smith, Muhammed Razzak 等NeurIPS 2025 · 被引用 11 次
- Active Statistical InferenceTijana Zrnic, Emmanuel J. CandèsICML 2024 · 被引用 34 次
- Efficient Active Learning for Gaussian Process Classification by Error ReductionGuang Zhao, Edward R. Dougherty, Byung-Jun Yoon, Francis J. Alexander 等NeurIPS 2021 · 被引用 28 次
- Learning -Statistics with Active InferenceXiaoning Wang, Huo Yuyang, Liuhua Peng, Changliang ZouICML 2026
