Pre-train or Annotate? Domain Adaptation with a Constrained Budget
Fan Bai, Alan Ritter, Wei Xu
摘要
Recent work has demonstrated that pretraining in-domain language models can boost performance when adapting to a new domain. However, the costs associated with pretraining raise an important question: given a fixed budget, what steps should an NLP practitioner take to maximize performance? In this paper, we view domain adaptation with a constrained budget as a consumer choice problem, where the goal is to select an optimal combination of data annotation and pre-training. We measure annotation costs of three procedural text datasets, along with the pre-training costs of several in-domain language models. The utility of different combinations of pretraining and data annotation are evaluated under varying budget constraints to assess which combination strategy works best. We find that for small budgets, spending all funds on annotation leads to the best performance; once the budget becomes large enough, however, a combination of data annotation and in-domain pre-training yields better performance. Our experiments suggest task-specific data annotation should be part of an economical strategy when adapting an NLP model to a new domain. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- MosaicBERT: A Bidirectional Encoder Optimized for Fast PretrainingJacob P. Portes, Alexander Trott, Sam Havens, Daniel King 等NeurIPS 2023 · 被引用 46 次
- Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is OffensiveTharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri 等EMNLP 2023 · 被引用 12 次
- Distill or Annotate? Cost-Efficient Fine-Tuning of Compact ModelsJunmo Kang, Wei Xu, Alan RitterACL 2023 · 被引用 5 次
- Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based FinetuningMohit Raghavendra, Junmo Kang, Alan RitterACL 2025 · 被引用 5 次
- LLMs are Better Than You Think: Label-Guided In-Context Learning for Named Entity RecognitionFan Bai, Hamid Hassanzadeh, Ardavan Saeedi, Mark DredzeEMNLP 2025 · 被引用 2 次
它引用的顶会 Paper3
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Transformer Based Multi-Source Domain AdaptationDustin Wright, Isabelle AugensteinEMNLP 2020 · 被引用 3 次
相关 Paper
- Adapting a Language Model While Preserving its General KnowledgeZixuan Ke, Yijia Shao, Haowei Lin, Hu Xu 等EMNLP 2022 · 被引用 6 次
- Scaling Laws for Forgetting during Finetuning with Pretraining Data InjectionLouis Béthune, David Grangier, Dan Busbridge, Eleonora Gualdoni 等ICML 2025
- Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference ModelsZachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion 等ICLR 2025 · 被引用 4 次
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 被引用 128 次
- Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models' MemoriesShizhe Diao, Tianyang Xu, Ruijia Xu, Jiawei Wang 等ACL 2023 · 被引用 17 次
