Lifelong Sequence Generation with Dynamic Module Expansion and Adaptation
Chengwei Qin, Chen Chen, Shafiq Joty
摘要
Lifelong sequence generation (LSG), a problem in continual learning, aims to continually train a model on a sequence of generation tasks to learn constantly emerging new generation patterns while avoiding the forgetting of previous knowledge. Existing LSG methods mainly focus on maintaining old knowledge while paying little attention to knowledge transfer across tasks. In contrast, humans can better learn new tasks by leveraging previously acquired knowledge from similar tasks. Inspired by the learning paradigm of humans, we propose Dynamic Module Expansion and Adaptation (DMEA), which enables the model to dynamically determine the architecture for acquiring new knowledge based on task correlation and select the most similar previous tasks to facilitate adaptation to new tasks. In addition, as the learning process can easily be biased towards the current task which might cause more severe forgetting of previously learned knowledge, we propose dynamic gradient scaling to balance the learning of the current task and replayed tasks. With extensive experiments, we demonstrate that DMEA can consistently outperform existing methods in different LSG settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- AnaCP: Toward Upper-Bound Continual Learning via Analytic Contrastive ProjectionSaleh Momeni, Changnan Xiao, Bing LiuNeurIPS 2025 · 被引用 8 次
- Beyond Output Matching: Bidirectional Alignment for Enhanced In-Context LearningChengwei Qin, Wenhan Xia, Fangkai Jiao, Chen Chen 等ACL 2025 · 被引用 7 次
- Grow-on-Demand: Sparse and Adaptive Expert Expansion for Continual Instruction TuningYing Zhang, Xingyue Guo, Yu Zhao, Xuhui Sui 等AAAI 2026
- Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot ManipulationYuanqi Yao, Siao Liu, Haoming Song, Delin Qu 等CVPR 2025
- SAPT: A Shared Attention Framework for Parameter-Efficient Continual Learning of Large Language ModelsWeixiang Zhao, Shilong Wang, Yulin Hu, Yanyan Zhao 等ACL 2024
它引用的顶会 Paper10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 被引用 247 次
- Continual Learning of a Mixed Sequence of Similar and Dissimilar TasksZixuan Ke, Bing Liu, Xingchang HuangNeurIPS 2020 · 被引用 173 次
- Achieving Forgetting Prevention and Knowledge Transfer in Continual LearningZixuan Ke, Bing Liu, Nianzu Ma, Hu Xu 等NeurIPS 2021 · 被引用 167 次
- End-to-End Neural Pipeline for Goal-Oriented Dialogue Systems using GPT-2DongHoon Ham, Jeong-Gwan Lee, Youngsoo Jang, Kee-Eung KimACL 2020 · 被引用 167 次
相关 Paper
- Continual Sequence Generation with Adaptive Compositional ModulesYanzhe Zhang, Xuezhi Wang, Diyi YangACL 2022 · 被引用 53 次
- Self-Expansion of Pre-trained Models with Mixture of Adapters for Continual LearningHuiyi Wang, Haodong Lu, Lina Yao, Dong GongCVPR 2025
- Efficient Continual Learning with Modular Networks and Task-Driven PriorsTom Veniat, Ludovic Denoyer, Marc'Aurelio RanzatoICLR 2021 · 被引用 110 次
- A Combinatorial Perspective on Transfer LearningJianan Wang, Eren Sezener, David Budden, Marcus Hutter 等NeurIPS 2020 · 被引用 9 次
- Self-Evolved Dynamic Expansion Model for Task-Free Continual LearningFei Ye, Adrian G. BorsICCV 2023 · 被引用 28 次
