On Statistical Bias In Active Learning: How and When to Fix It
Sebastian Farquhar, Yarin Gal, Tom Rainforth
摘要
Active learning is a powerful tool when labelling data is expensive, but it introduces a bias because the training data no longer follows the population distribution. We formalize this bias and investigate the situations in which it can be harmful and sometimes even helpful. We further introduce novel corrective weights to remove bias when doing so is beneficial. Through this, our work not only provides a useful mechanism that can improve the active learning approach, but also an explanation of the empirical successes of various existing approaches which ignore this bias. In particular, we show that this bias can be actively helpful when training overparameterized models -- like neural networks -- with relatively little data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntSören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma 等ICML 2022 · 被引用 237 次
- Active Testing: Sample-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Tom RainforthICML 2021 · 被引用 81 次
- Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational DataAndrew Jesson, Panagiotis Tigas, Joost van Amersfoort, Andreas Kirsch 等NeurIPS 2021 · 被引用 42 次
- Active Surrogate Estimators: An Active Learning Approach to Label-Efficient Model EvaluationJannik Kossen, Sebastian Farquhar, Yarin Gal, Thomas RainforthNeurIPS 2022 · 被引用 36 次
- A Lagrangian Duality Approach to Active LearningJuan Elenter, Navid NaderiAlizadeh, Alejandro RibeiroNeurIPS 2022 · 被引用 31 次
它引用的顶会 Paper3
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford 等ICLR 2020 · 被引用 974 次
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 被引用 662 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
相关 Paper
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch 等EMNLP 2020 · 被引用 144 次
- Is Importance Weighting Incompatible with Interpolating Classifiers?Ke Alexander Wang, Niladri Shekhar Chatterji, Saminul Haque, Tatsunori HashimotoICLR 2022 · 被引用 22 次
- Towards Balanced Active Learning for Multimodal ClassificationMeng Shen, Yizheng Huang, Jianxiong Yin, Heqing Zou 等ACM MM 2023 · 被引用 5 次
- Understanding Instance-Level Label Noise: Disparate Impacts and TreatmentsYang LiuICML 2021 · 被引用 38 次
- Thumb on the Scale: Optimal Loss Weighting in Last Layer RetrainingNathan Stromberg, Christos Thrampoulidis, Lalitha SankarNeurIPS 2025 · 被引用 1 次
