Batch Active Learning at Scale
Gui Citovsky, Giulia DeSalvo, Claudio Gentile, Lazaros Karydas, Anand Rajagopalan, Afshin Rostamizadeh, Sanjiv Kumar
摘要
The ability to train complex and highly effective models often requires an abundance of training data, which can easily become a bottleneck in cost, time, and computational resources. Batch active learning, which adaptively issues batched queries to a labeling oracle, is a common approach for addressing this problem. The practical benefits of batch sampling come with the downside of less adaptivity and the risk of sampling redundant examples within a batch -- a risk that grows with the batch size. In this work, we analyze an efficient active learning algorithm, which focuses on the large batch setting. In particular, we show that our sampling method, which combines notions of uncertainty and diversity, easily scales to batch sizes (100K-1M) several orders of magnitude larger than used in previous studies and provides significant improvements in model training efficiency compared to recent baselines. Finally, we provide an initial theoretical analysis, proving label complexity guarantees for a related sampling method, which we show is approximately equivalent to our sampling method in specific settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- A Survey of Active Learning for Natural Language ProcessingZhisong Zhang, Emma Strubell, Eduard H. HovyEMNLP 2022 · 被引用 60 次
- GALAXY: Graph-based Active Learning at the ExtremeJifan Zhang, Julian Katz-Samuels, Robert D. NowakICML 2022 · 被引用 47 次
- Robust Data Pruning under Label Noise via Maximizing Re-labeling AccuracyDongmin Park, Seola Choi, Doyoung Kim, Hwanjun Song 等NeurIPS 2023 · 被引用 42 次
- Algorithm Selection for Deep Active Learning with Imbalanced DatasetsJifan Zhang, Shuai Shao, Saurabh Verma, Robert D. NowakNeurIPS 2023 · 被引用 38 次
- Meta-Query-Net: Resolving Purity-Informativeness Dilemma in Open-set Active LearningDongmin Park, Yooju Shin, Jihwan Bang, Youngjun Lee 等NeurIPS 2022 · 被引用 37 次
它引用的顶会 Paper3
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford 等ICLR 2020 · 被引用 974 次
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 被引用 662 次
- Task-Aware Variational Adversarial Active LearningKwanyoung Kim, Dongwon Park, Kwang In Kim, Se Young ChunCVPR 2021
相关 Paper
- Scalable Batch-Mode Deep Bayesian Active Learning via Equivalence Class AnnealingRenyu Zhang, Aly A. Khan, Robert L. Grossman, Yuxin ChenICLR 2023 · 被引用 1 次
- Achieving Minimax Rates in Pool-Based Batch Active LearningClaudio Gentile, Zhilei Wang, Tong ZhangICML 2022 · 被引用 16 次
- Active Learning for Multiple Target ModelsYing-Peng Tang, Sheng-Jun HuangNeurIPS 2022 · 被引用 7 次
- Provably Neural Active Learning Succeeds via Prioritizing Perplexing SamplesDake Bu, Wei Huang, Taiji Suzuki, Ji Cheng 等ICML 2024 · 被引用 5 次
- Asking the Right Questions to the Right Users: Active Learning with Imperfect OraclesShayok ChakrabortyAAAI 2020 · 被引用 23 次
