Pre-train or Annotate? Domain Adaptation with a Constrained Budget
Fan Bai, Alan Ritter, Wei Xu
Abstract
Recent work has demonstrated that pretraining in-domain language models can boost performance when adapting to a new domain. However, the costs associated with pretraining raise an important question: given a fixed budget, what steps should an NLP practitioner take to maximize performance? In this paper, we view domain adaptation with a constrained budget as a consumer choice problem, where the goal is to select an optimal combination of data annotation and pre-training. We measure annotation costs of three procedural text datasets, along with the pre-training costs of several in-domain language models. The utility of different combinations of pretraining and data annotation are evaluated under varying budget constraints to assess which combination strategy works best. We find that for small budgets, spending all funds on annotation leads to the best performance; once the budget becomes large enough, however, a combination of data annotation and in-domain pre-training yields better performance. Our experiments suggest task-specific data annotation should be part of an economical strategy when adapting an NLP model to a new domain. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 314b099a-1852-4fe3-a79f-f3f0e06b99b4Cited by top-tier papers7
- MosaicBERT: A Bidirectional Encoder Optimized for Fast PretrainingJacob P. Portes, Alexander Trott, Sam Havens, Daniel King et al.NeurIPS 2023 · 46 citations
- Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is OffensiveTharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri et al.EMNLP 2023 · 12 citations
- Distill or Annotate? Cost-Efficient Fine-Tuning of Compact ModelsJunmo Kang, Wei Xu, Alan RitterACL 2023 · 5 citations
- Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based FinetuningMohit Raghavendra, Junmo Kang, Alan RitterACL 2025 · 5 citations
- LLMs are Better Than You Think: Label-Guided In-Context Learning for Named Entity RecognitionFan Bai, Hamid Hassanzadeh, Ardavan Saeedi, Mark DredzeEMNLP 2025 · 2 citations
Builds on3
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 340 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Transformer Based Multi-Source Domain AdaptationDustin Wright, Isabelle AugensteinEMNLP 2020 · 3 citations
Related papers
- Adapting a Language Model While Preserving its General KnowledgeZixuan Ke, Yijia Shao, Haowei Lin, Hu Xu et al.EMNLP 2022 · 6 citations
- Scaling Laws for Forgetting during Finetuning with Pretraining Data InjectionLouis Béthune, David Grangier, Dan Busbridge, Eleonora Gualdoni et al.ICML 2025
- Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference ModelsZachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion et al.ICLR 2025 · 4 citations
- Cold-start Active Learning through Self-supervised Language ModelingMichelle Yuan, Hsuan-Tien Lin, Jordan L. Boyd-GraberEMNLP 2020 · 128 citations
- Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models' MemoriesShizhe Diao, Tianyang Xu, Ruijia Xu, Jiawei Wang et al.ACL 2023 · 17 citations
