Recasting Continual Learning as Sequence Modeling
Soochan Lee, Jaehyeon Son, Gunhee Kim
Abstract
In this work, we aim to establish a strong connection between two significant bodies of machine learning research: continual learning and sequence modeling. That is, we propose to formulate continual learning as a sequence modeling problem, allowing advanced sequence models to be utilized for continual learning. Under this formulation, the continual learning process becomes the forward pass of a sequence model. By adopting the meta-continual learning (MCL) framework, we can train the sequence model at the meta-level, on multiple continual learning episodes. As a specific example of our new formulation, we demonstrate the application of Transformers and their efficient variants as MCL methods. Our experiments on seven benchmarks, covering both classification and regression, show that sequence models can be an attractive solution for general MCL. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4835d75-cf3a-4de6-9f07-1313cff75d8bCited by top-tier papers4
- Learning to Continually Learn with the Bayesian PrincipleSoochan Lee, Hyeonseong Jeon, Jaehyeon Son, Gunhee KimICML 2024 · 11 citations
- Compositional-ARC: Assessing Systematic Generalization in Abstract Spatial ReasoningPhilipp Mondorf, Shijia Zhou, Monica Riedler, Barbara PlankICLR 2026 · 3 citations
- In-context Learning of Evolving Data Streams with Tabular Foundational ModelsAfonso Lourenço, João Gama, Eric P. Xing, Goreti MarreirosKDD 2026 · 2 citations
- Distilling Reinforcement Learning Algorithms for In-Context Model-Based PlanningJaehyeon Son, Soochan Lee, Gunhee KimICLR 2025
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Transformer protein language models are unsupervised structure learnersRoshan Rao, Joshua Meier, Tom Sercu, Sergey Ovchinnikov et al.ICLR 2021 · 366 citations
- LAMOL: LAnguage MOdeling for Lifelong Language LearningFan-Keng Sun, Cheng-Hao Ho, Hung-Yi LeeICLR 2020 · 247 citations
Related papers
- Compositional Language Continual LearningYuanpeng Li, Liang Zhao, Kenneth Church, Mohamed ElhoseinyICLR 2020 · 40 citations
- Continual Sequence Generation with Adaptive Compositional ModulesYanzhe Zhang, Xuezhi Wang, Diyi YangACL 2022 · 53 citations
- A Simple Recipe to Meta-Learn Forward and Backward TransferEdoardo Cetin, Antonio Carta, Oya ÇeliktutanICCV 2023 · 1 citation
- Meta-attention for ViT-backed Continual LearningMengqi Xue, Haofei Zhang, Jie Song, Mingli SongCVPR 2022 · 40 citations
- Continuous Meta-Learning without TasksJames Harrison, Apoorva Sharma, Chelsea Finn, Marco PavoneNeurIPS 2020 · 86 citations
