LOIRE: LifelOng learning on Incremental data via pre-trained language model gRowth Efficiently
Xue Han, Yitong Wang, Junlan Feng, Wenchun Gao, Qian Hu, Chao Deng
摘要
Large-scale pre-trained language models (PLMs) require significant computational resources to train from scratch on large volumes of data. But in the real world, emerging data from diverse sources may not be initially available for pretraining. Recent studies on lifelong learning have tried to solve this problem by exploring the use of model growth techniques to effectively incorporate new knowledge without the need for complete re-training. However, model growth approaches utilized have issues with growth operators that do not ensure strict function preservation or growth schedules that only include a few growth dimensions, reducing lifelong learning's effect. Furthermore, existing approaches often assume that emerging data has the same distribution as pre-training data, causing catastrophic forgetting of previously acquired knowledge. To address the aforementioned issues, we introduce LOIRE, a framework for lifelong learning that enables PLMs to effectively grow their capacity using incremental data. LOIRE employs growth operators for all feasible dimensions and a growth schedule to generate the optimal expansion sequence in the field of lifelong learning. Specifically, we present a novel plug-in layer growth operator with residual connections that skip the newly added layer during initial training while ensuring function preservation. We additionally propose an iterative distillation strategy for LOIRE that allows an intermediate model in the growth stages to switch between being a student and a teacher, reducing catastrophic forgetting during growth. Experiments show that LOIRE can reduce computational expenses by an average of 29.22% while retaining equivalent or better downstream performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive LearningQifan Yu, Xinyu Ma, Zhijian Zhuo, Minrui Wang 等ICML 2026 · 被引用 3 次
- Escaping the Subspace Trap: The Role of Optimizer Geometry in Model Width ExpansionJiabei Chen, Haoyu Wang, Yang Yu, Yao Xu 等ICML 2026
- R-Tuning: Wavelet-Decomposed Replay and Semantic Alignment for Continual Adaptation of Pretrained Time-Series ModelsTianyi Yin, Jingwei Wang, Chenze Wang, Han Wang 等AAAI 2026
- Soft Orthogonal Low-Rank Adaptation for Knowledge Sharing in Large Language Model Continual LearningYitong Wang, Xue Han, Wenchun Gao, Qian Hu 等ACL 2026
- ORTCL: Towards Continual Learning of Time Series Foundation Models on Streaming Data via Orthogonal RotationLi Lin, Xinrui Zhang, Qi Zhang, Shuai Wang 等AAAI 2026
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Post-Training Quantization for Vision TransformerZhenhua Liu, Yunhe Wang, Kai Han, Wei Zhang 等NeurIPS 2021 · 被引用 528 次
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney 等ACL 2020 · 被引用 424 次
- Language Modeling Is CompressionGrégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt 等ICLR 2024 · 被引用 243 次
相关 Paper
- Lifelong Language Pretraining with Distribution-Specialized ExpertsWuyang Chen, Yanqi Zhou, Nan Du, Yanping Huang 等ICML 2023 · 被引用 85 次
- Grow-on-Demand: Sparse and Adaptive Expert Expansion for Continual Instruction TuningYing Zhang, Xingyue Guo, Yu Zhao, Xuhui Sui 等AAAI 2026
- Pretrained Language Model in Continual Learning: A Comparative StudyTongtong Wu, Massimo Caccia, Zhuang Li, Yuan-Fang Li 等ICLR 2022 · 被引用 76 次
- Self-Regulating Prompt Expansion for Continual LearningYiwen Wang, Diana Benavides-Prado, Yun Sing KohKDD 2026
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 被引用 247 次
