Lifelong Language Knowledge Distillation
Yung-Sung Chuang, Shang-Yu Su, Yun-Nung Chen
摘要
It is challenging to perform lifelong language learning (LLL) on a stream of different tasks without any performance degradation comparing to the multi-task counterparts. To address this issue, we present Lifelong Language Knowledge Distillation (L2KD), a simple but efficient method that can be easily applied to existing LLL architectures in order to mitigate the degradation. Specifically, when the LLL model is trained on a new task, we assign a teacher model to first learn the new task, and pass the knowledge to the LLL model via knowledge distillation. Therefore, the LLL model can better adapt to the new task while keeping the previously learned knowledge. Experiments show that the proposed L2KD consistently improves previous state-ofthe-art models, and the degradation comparing to multi-task models in LLL tasks is well mitigated for both sequence generation and text classification tasks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Achieving Forgetting Prevention and Knowledge Transfer in Continual LearningZixuan Ke, Bing Liu, Nianzu Ma, Hu Xu 等NeurIPS 2021 · 被引用 167 次
- LFPT5: A Unified Framework for Lifelong Few-shot Language Learning Based on Prompt Tuning of T5Chengwei Qin, Shafiq R. JotyICLR 2022 · 被引用 128 次
- Memory Efficient Continual Learning with TransformersBeyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal 等NeurIPS 2022 · 被引用 75 次
- Continual Learning in Task-Oriented Dialogue SystemsAndrea Madotto, Zhaojiang Lin, Zhenpeng Zhou, Seungwhan Moon 等EMNLP 2021 · 被引用 68 次
- Continual Sequence Generation with Adaptive Compositional ModulesYanzhe Zhang, Xuezhi Wang, Diyi YangACL 2022 · 被引用 53 次
它引用的顶会 Paper3
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 被引用 247 次
- Understanding Knowledge Distillation in Non-autoregressive Machine TranslationChunting Zhou, Jiatao Gu, Graham NeubigICLR 2020 · 被引用 235 次
- Lifelong GAN: Continual Learning for Conditional Image GenerationMengyao Zhai, Lei Chen, Frederick Tung, Jiawei He 等ICCV 2019 · 被引用 204 次
相关 Paper
- Revisiting Knowledge Distillation for Autoregressive Language ModelsQihuang Zhong, Liang Ding, Li Shen, Juhua Liu 等ACL 2024
- Context Distillation Retains Post-Training Capabilities in Continually Trained LMsShankar Padmanabhan, Mustafa Omer Gul, Tanya GoyalICML 2026
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang 等NeurIPS 2024 · 被引用 50 次
- Text Representation Distillation via Information Bottleneck PrincipleYanzhao Zhang, Dingkun Long, Zehan Li, Pengjun XieEMNLP 2023 · 被引用 4 次
- AdaEdit: Advancing Continuous Knowledge Editing For Large Language ModelsQi Li, Xiaowen ChuACL 2025
