Adapting a Language Model While Preserving its General Knowledge
Zixuan Ke, Yijia Shao, Haowei Lin, Hu Xu, Lei Shu, Bing Liu
摘要
Domain-adaptive pre-training (or DA-training for short), also known as post-training, aims to train a pre-trained general-purpose language model (LM) using an unlabeled corpus of a particular domain to adapt the LM so that endtasks in the domain can give improved performances. However, existing DA-training methods are in some sense blind as they do not explicitly identify what knowledge in the LM should be preserved and what should be changed by the domain corpus. This paper shows that the existing methods are suboptimal and proposes a novel method to perform a more informed adaptation of the knowledge in the LM by (1) soft-masking the attention heads based on their importance to best preserve the general knowledge in the LM and (2) contrasting the representations of the general and the full (both general and domain knowledge) to learn an integrated representation with both general and domain-specific knowledge. Experimental results will demonstrate the effectiveness of the proposed approach. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Selecting Large Language Model to Fine-tune via Rectified Scaling LawHaowei Lin, Baizhou Huang, Haotian Ye, Qinyu Chen 等ICML 2024 · 被引用 32 次
- Mixture-of-Domain-Adapters: Decoupling and Injecting Domain Knowledge to Pre-trained Language Models' MemoriesShizhe Diao, Tianyang Xu, Ruijia Xu, Jiawei Wang 等ACL 2023 · 被引用 17 次
- Continual Pre-training of Language ModelsZixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi 等ICLR 2023 · 被引用 15 次
- Once Read is Enough: Domain-specific Pretraining-free Language Models with Cluster-guided Sparse Experts for Long-tail Domain KnowledgeFang Dong, Mengyi Chen, Jixian Zhou, Yubin Shi 等NeurIPS 2024 · 被引用 6 次
- Demystifying Domain-adaptive Post-training for Financial LLMsZixuan Ke, Yifei Ming, Xuan-Phi Nguyen, Caiming Xiong 等EMNLP 2025 · 被引用 2 次
它引用的顶会 Paper12
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 被引用 695 次
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu 等NeurIPS 2020 · 被引用 428 次
相关 Paper
- Adversarial and Domain-Aware BERT for Cross-Domain Sentiment AnalysisChunning Du, Haifeng Sun, Jingyu Wang, Qi Qi 等ACL 2020 · 被引用 165 次
- ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled TuningJinyang Zhang, Yue Fang, Hongxin Ding, Weibin Liao 等ICLR 2026 · 被引用 5 次
- LLaMA Pro: Progressive LLaMA with Block ExpansionChengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu 等ACL 2024
- Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model AdaptationMinki Kang, Moonsu Han, Sung Ju HwangEMNLP 2020 · 被引用 12 次
- G-MAP: General Memory-Augmented Pre-trained Language Model for Domain TasksZhongwei Wan, Yichun Yin, Wei Zhang, Jiaxin Shi 等EMNLP 2022 · 被引用 2 次
