On the Fragility of Active Learners for Text Classification
Abhishek Ghose, Emma Nguyen
摘要
Active learning (AL) techniques optimally utilize a labeling budget by iteratively selecting instances that are most valuable for learning. However, they lack "prerequisite checks", i.e., there are no prescribed criteria to pick an AL algorithm best suited for a dataset. A practitioner must pick a technique they trust would beat random sampling, based on prior reported results, and hope that it is resilient to the many variables in their environment: dataset, labeling budget and prediction pipelines. The important questions then are: how often on average, do we expect any AL technique to reliably beat the computationally cheap and easy-to-implement strategy of random sampling? Does it at least make sense to use AL in an "Always ON" mode in a prediction pipeline, so that while it might not always help, it never under-performs random sampling? How much of a role does the prediction pipeline play in AL's success? We examine these questions in detail for the task of text classification using pre-trained representations, which are ubiquitous today. Our primary contribution here is a rigorous evaluation of AL techniques, old and new, across setups that vary wrt datasets, text representations and classifiers. This unlocks multiple insights around warm-up times, i.e., number of labels before gains from AL are seen, viability of an "Always ON" mode and the relative significance of different factors. Additionally, we release a framework for rigorous benchmarking of AL techniques for text classification.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch 等EMNLP 2020 · 被引用 144 次
- Navigating the Pitfalls of Active Learning Evaluation: A Systematic Framework for Meaningful Performance AssessmentCarsten T. Lüth, Till J. Bungert, Lukas Klein, Paul F. JaegerNeurIPS 2023 · 被引用 32 次
- Active Learning by Acquiring Contrastive ExamplesKaterina Margatina, Giorgos Vernikos, Loïc Barrault, Nikolaos AletrasEMNLP 2021 · 被引用 8 次
相关 Paper
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 被引用 128 次
- CoLAL: Co-learning Active Learning for Text ClassificationLinh Le, Genghong Zhao, Xia Zhang, Guido Zuccon 等AAAI 2024 · 被引用 5 次
- Active Multi-Task Representation LearningYifang Chen, Kevin Jamieson, Simon S. DuICML 2022 · 被引用 18 次
- Cleaning the Pool: Progressive Filtering of Unlabeled Pools in Deep Active LearningDenis Huseljic, Marek Herde, Lukas Rauch, Paul Hahn 等CVPR 2026 · 被引用 2 次
- Active Learning Helps Pretrained Models Learn the Intended TaskAlex Tamkin, Dat Nguyen, Salil Deshpande, Jesse Mu 等NeurIPS 2022 · 被引用 54 次
