Mitigating Forgetting in Adapting Pre-trained Language Models to Text Processing Tasks via Consistency Alignment
Jianqi Gao, Hao Wu, Yiu-ming Cheung, Jian Cao, Hang Yu, Yonggang Zhang
Abstract
There are a large number of text processing tasks in web applications, such as sentiment classification, summary extraction, and question answering. Recently, fine-tuning pre-trained language models (PLMs) to adapt to downstream text-processing tasks has attracted much attention. However, due to the differences in data, model, and tasks between the pre-training and fine-tuning processes, the fine-tuning process may suffer from catastrophic forgetting of pre-training knowledge, which may implicitly limit the model's performance and generalization ability. To address these challenges, we propose a novel dual-model framework, termed as consistency alignment (CoAi). The insight of CoAi lies in building an auxiliary model that simulates the distribution of pre-training knowledge in real-time according to the current task, and co-training the task-specific model and the auxiliary model to balance the pre-training knowledge and task-specific knowledge during fine-tuning. Specifically, the auxiliary model is constructed on-the-fly to maintain the pre-training knowledge. Subsequently, CoAi simulates the pre-training process by performing distributional exploration in the parameter space, which is built upon our novel insight into the transformation between data and model parameter space. However, the objectives leveraged to construct the auxiliary model lead to the misalignment between the pre-training and task-specific knowledge. To alleviate the inconsistency, we employ an auxiliary variable to align the prediction distribution of the task-specific and the auxiliary models, inspired by constrastive clustering. We validate the effectiveness of CoAi on nine classic classification tasks and three generation tasks, showing consistent and significant improvements compared with state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 2f1e90ed-ba13-4376-a5e6-30823478c276Cited by top-tier papers2
- Breaking Down Market Barriers: Distilled Prompt-Tuning Approach for Cross-Market RecommendationLeqi Zhang, Wayne Lu, Haiyang Zhang, Elliott Wen et al.AAAI 2026
- An Effective Levelling Paradigm for Unlabeled ScenariosFangming Cui, Zhou Yu, Di Yang, Yuqiang Ren et al.NeurIPS 2025
Related papers
- Preserving Commonsense Knowledge from Pre-trained Language Models via Causal InferenceJunhao Zheng, Qianli Ma, Shengjie Qiu, Yue Wu et al.ACL 2023 · 9 citations
- Mix-CPT: A Domain Adaptation Framework via Decoupling Knowledge Learning and Format AlignmentJinhao Jiang, Junyi Li, Xin Zhao, Yang Song et al.ICLR 2025
- Adapt Before Continual LearningAojun Lu, Tao Feng, Hangjie Yuan, Chunhui Ding et al.AAAI 2026
- G-MAP: General Memory-Augmented Pre-trained Language Model for Domain TasksZhongwei Wan, Yichun Yin, Wei Zhang, Jiaxin Shi et al.EMNLP 2022 · 2 citations
- From Dense to Sparse: Contrastive Pruning for Better Pre-trained Language Model CompressionRunxin Xu, Fuli Luo, Chengyu Wang, Baobao Chang et al.AAAI 2022 · 32 citations
