An Empirical Investigation Towards Efficient Multi-Domain Language Model Pre-training
Kristjan Arumae, Qing Sun, Parminder Bhatia
Abstract
Pre-training large language models has become a standard in the natural language processing community. Such models are pretrained on generic data (e.g. BookCorpus and English Wikipedia) and often fine-tuned on tasks in the same domain. However, in order to achieve state-of-the-art performance on out of domain tasks such as clinical named entity recognition and relation extraction, additional in domain pre-training is required. In practice, staged multi-domain pre-training presents performance deterioration in the form of catastrophic forgetting (CF) when evaluated on a generic benchmark such as GLUE. In this paper we conduct an empirical investigation into known methods to mitigate CF. We find that elastic weight consolidation provides best overall scores yielding only a 0.33% drop in performance across seven generic tasks while remaining competitive in bio-medical tasks. Furthermore, we explore gradient and latent clustering based data selection techniques to improve coverage when using elastic weight consolidation and experience replay methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on3
- TANDA: Transfer and Adapt Pre-Trained Transformer Models for Answer Sentence SelectionSiddhant Garg, Thuy Vu, Alessandro MoschittiAAAI 2020 · 229 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Unsupervised Domain Clusters in Pretrained Language ModelsRoee Aharoni, Yoav GoldbergACL 2020 · 13 citations
Related papers
- Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to AlgorithmHaoyu Wang, yifan shang, Zhongxiang Sun, Weijie Yu et al.ICML 2026
- G-MAP: General Memory-Augmented Pre-trained Language Model for Domain TasksZhongwei Wan, Yichun Yin, Wei Zhang, Jiaxin Shi et al.EMNLP 2022 · 2 citations
- Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less ForgettingSanyuan Chen, Yutai Hou, Yiming Cui, Wanxiang Che et al.EMNLP 2020 · 152 citations
- Forget Forgetting: Continual Learning in a World of Abundant MemoryDongkyu Cho, Taesup Moon, Rumi Chunara, Kyunghyun Cho et al.ICLR 2026 · 9 citations
- Prototypical Replay with Old-class Focusing Knowledge Distillation for Incremental Named Entity RecognitionZesheng Liu, Qiannan Zhu, Cuiping Li, Hong ChenAAAI 2025
