Continual Pre-training of Language Models
Zixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi, Gyuhak Kim, Bing Liu
摘要
Language models (LMs) have been instrumental for the rapid advance of natural language processing. This paper studies continual pre-training of LMs, in particular, continual domain-adaptive pre-training (or continual DAP-training). Existing research has shown that further pre-training an LM using a domain corpus to adapt the LM to the domain can improve the end-task performance in the domain. This paper proposes a novel method to continually DAP-train an LM with a sequence of unlabeled domain corpora to adapt the LM to these domains to improve their endtask performances. The key novelty of our method is a soft-masking mechanism that directly controls the update to the LM. A novel proxy is also proposed to preserve the general knowledge in the original LM. Additionally, it contrasts the representations of the previously learned domain knowledge (including the general knowledge in the pre-trained LM) and the knowledge from the current full network to achieve knowledge integration. The method not only overcomes catastrophic forgetting, but also achieves knowledge transfer to improve end-task performances. Empirical evaluation demonstrates the effectiveness of the proposed method. 1 * Equal contribution † The work was done when this author was visiting Bing Liu's group at University of Illinois at Chicago.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper61
- Hierarchical Decomposition of Prompt-Based Continual Learning: Rethinking Obscured Sub-optimalityLiyuan Wang, Jingyi Xie, Xingxing Zhang, Mingyi Huang 等NeurIPS 2023 · 被引用 183 次
- Learning to Discover at Test TimeMert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi 等ICML 2026 · 被引用 73 次
- Functional Interpolation for Relative Positions improves Long Context TransformersShanda Li, Chong You, Guru Guruganesh, Joshua Ainslie 等ICLR 2024 · 被引用 66 次
- Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language ModelsShangbin Feng, Weijia Shi, Yuyang Bai, Vidhisha Balachandran 等ICLR 2024 · 被引用 56 次
- D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language ModelsHaoran Que, Jiaheng Liu, Ge Zhang, Chenchen Zhang 等NeurIPS 2024 · 被引用 47 次
它引用的顶会 Paper20
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn 等ICLR 2022 · 被引用 527 次
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney 等ACL 2020 · 被引用 424 次
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi 等NeurIPS 2020 · 被引用 364 次
- Towards Continual Knowledge Learning of Language ModelsJoel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin 等ICLR 2022 · 被引用 204 次
相关 Paper
- Adapting a Language Model While Preserving its General KnowledgeZixuan Ke, Yijia Shao, Haowei Lin, Hu Xu 等EMNLP 2022 · 被引用 6 次
- G-MAP: General Memory-Augmented Pre-trained Language Model for Domain TasksZhongwei Wan, Yichun Yin, Wei Zhang, Jiaxin Shi 等EMNLP 2022 · 被引用 2 次
- Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format AlignmentJinhao Jiang, Junyi Li, Xin Zhao, Yang Song 等ICLR 2025
- CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language ModelsJiawei Gu, Zacc Yang, Chuanghao Ding, Rui Zhao 等EMNLP 2024 · 被引用 2 次
- ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled TuningJinyang Zhang, Yue Fang, Hongxin Ding, Weibin Liao 等ICLR 2026 · 被引用 5 次
