Cold-start Active Learning through Self-supervised Language Modeling
Michelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-Graber
Abstract
Active learning strives to reduce annotation costs by choosing the most critical examples to label. Typically, the active learning strategy is contingent on the classification model. For instance, uncertainty sampling depends on poorly calibrated model confidence scores. In the cold-start setting, active learning is impractical because of model instability and data scarcity. Fortunately, modern NLP provides an additional source of information: pretrained language models. The pre-training loss can find examples that surprise the model and should be labeled for efficient fine-tuning. Therefore, we treat the language modeling loss as a proxy for classification uncertainty. With BERT, we develop a simple strategy based on the masked language modeling loss that minimizes labeling costs for text classification. Compared to other baselines, our approach reaches higher accuracy within less sampling iterations and computation time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e1a1fb4-0dfd-4a3a-86d4-295ac4d4a3ceCited by top-tier papers35
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 383 citations
- Active Learning on a Budget: Opposite Strategies Suit High and Low BudgetsGuy Hacohen, Avihu Dekel, Daphna WeinshallICML 2022 · 163 citations
- Active Learning Through a Covering LensOfer Yehuda, Avihu Dekel, Guy Hacohen, Daphna WeinshallNeurIPS 2022 · 102 citations
- A Survey of Active Learning for Natural Language ProcessingZhisong Zhang, Emma Strubell, Eduard H. HovyEMNLP 2022 · 60 citations
- Active Learning Helps Pretrained Models Learn the Intended TaskAlex Tamkin, Dat Nguyen, Salil Deshpande, Jesse Mu et al.NeurIPS 2022 · 54 citations
Builds on4
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- On Warm-Starting Neural Network TrainingJordan T. Ash, Ryan P. AdamsNeurIPS 2020 · 288 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
- Interactive Refinement of Cross-Lingual Word EmbeddingsMichelle Yuan, Mozhi Zhang, Benjamin Van Durme, Leah Findlater et al.EMNLP 2020 · 32 citations
Related papers
- Active Learning for BERT: An Empirical StudyLiat Ein-Dor, Alon Halfon, Ariel Gera, Eyal Shnarch et al.EMNLP 2020 · 144 citations
- Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language ModelsChristopher Schröder, Gerhard HeyerEMNLP 2024 · 2 citations
- On the Fragility of Active Learners for Text ClassificationAbhishek Ghose, Emma NguyenEMNLP 2024 · 2 citations
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 182 citations
- Active Bayesian Assessment of Black-Box ClassifiersDisi Ji, Robert L. Logan IV, Padhraic Smyth, Mark SteyversAAAI 2021 · 3 citations
