Lifelong Language Knowledge Distillation
Yung-Sung Chuang, Shang-Yu Su, Yun-Nung Chen
Abstract
It is challenging to perform lifelong language learning (LLL) on a stream of different tasks without any performance degradation comparing to the multi-task counterparts. To address this issue, we present Lifelong Language Knowledge Distillation (L2KD), a simple but efficient method that can be easily applied to existing LLL architectures in order to mitigate the degradation. Specifically, when the LLL model is trained on a new task, we assign a teacher model to first learn the new task, and pass the knowledge to the LLL model via knowledge distillation. Therefore, the LLL model can better adapt to the new task while keeping the previously learned knowledge. Experiments show that the proposed L2KD consistently improves previous state-ofthe-art models, and the degradation comparing to multi-task models in LLL tasks is well mitigated for both sequence generation and text classification tasks. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Achieving Forgetting Prevention and Knowledge Transfer in Continual LearningZixuan Ke, Bing Liu, Nianzu Ma, Hu Xu et al.NeurIPS 2021 · 167 citations
- LFPT5: A Unified Framework for Lifelong Few-shot Language Learning Based on Prompt Tuning of T5Chengwei Qin, Shafiq R. JotyICLR 2022 · 128 citations
- Memory Efficient Continual Learning with TransformersBeyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal et al.NeurIPS 2022 · 75 citations
- Continual Learning in Task-Oriented Dialogue SystemsAndrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon et al.EMNLP 2021 · 68 citations
- Continual Sequence Generation with Adaptive Compositional ModulesYanzhe Zhang, Xuezhi Wang, Diyi YangACL 2022 · 53 citations
Builds on3
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 247 citations
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 235 citations
- Lifelong GAN: Continual Learning for Conditional Image GenerationMengyao Zhai, Lei Chen, Frederick Tung, Jiawei He et al.ICCV 2019 · 204 citations
Related papers
- Revisiting Knowledge Distillation for Autoregressive Language ModelsQihuang Zhong, Liang Ding, Li Shen, Juhua Liu et al.ACL 2024
- Context Distillation Retains Post-Training Capabilities in Continually Trained LMsShankar Padmanabhan, Mustafa Omer Gul, Tanya GoyalICML 2026
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang et al.NeurIPS 2024 · 50 citations
- Text Representation Distillation via Information Bottleneck PrincipleYanzhao Zhang, Dingkun Long, Zehan Li, Pengjun XieEMNLP 2023 · 4 citations
- AdaEdit: Advancing Continuous Knowledge Editing For Large Language ModelsQi Li, Xiaowen ChuACL 2025
