On Statistical Bias In Active Learning: How and When to Fix It
Sebastian Farquhar, Yarin Gal, Tom Rainforth
Abstract
Active learning is a powerful tool when labelling data is expensive, but it introduces a bias because the training data no longer follows the population distribution. We formalize this bias and investigate the situations in which it can be harmful and sometimes even helpful. We further introduce novel corrective weights to remove bias when doing so is beneficial. Through this, our work not only provides a useful mechanism that can improve the active learning approach, but also an explanation of the empirical successes of various existing approaches which ignore this bias. In particular, we show that this bias can be actively helpful when training overparameterized models -- like neural networks -- with relatively little data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f212af69-07b7-4e48-a236-782d33bb1fe2Cited by top-tier papers22
- Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntSören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma et al.ICML 2022 · 237 citations
- Active Testing: Sample-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Tom RainforthICML 2021 · 81 citations
- Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational DataAndrew Jesson, Panagiotis Tigas, Joost van Amersfoort, Andreas Kirsch et al.NeurIPS 2021 · 42 citations
- Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Thomas RainforthNeurIPS 2022 · 36 citations
- A Lagrangian Duality Approach to Active LearningJuan Elenter, Navid NaderiAlizadeh, Alejandro RibeiroNeurIPS 2022 · 31 citations
Builds on3
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 662 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
Related papers
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch et al.EMNLP 2020 · 144 citations
- Is Importance Weighting Incompatible with Interpolating Classifiers?Ke Alexander Wang, Niladri Shekhar Chatterji, Saminul Haque, Tatsunori HashimotoICLR 2022 · 22 citations
- Towards Balanced Active Learning for Multimodal ClassificationMeng Shen, Yizheng Huang, Jianxiong Yin, Heqing Zou et al.ACM MM 2023 · 5 citations
- Understanding Instance-Level Label Noise: Disparate Impacts and TreatmentsYang LiuICML 2021 · 38 citations
- Thumb on the Scale: Optimal Loss Weighting in Last Layer RetrainingNathan Stromberg, Christos Thrampoulidis, Lalitha SankarNeurIPS 2025 · 1 citation
