Continual learning with hypernetworks
Johannes von Oswald, Christian Henning, João Sacramento, Benjamin F. Grewe
摘要
Artificial neural networks suffer from catastrophic forgetting when they are sequentially trained on multiple tasks. To overcome this problem, we present a novel approach based on task-conditioned hypernetworks, i.e., networks that generate the weights of a target model based on task identity. Continual learning (CL) is less difficult for this class of models thanks to a simple key feature: instead of recalling the input-output relations of all previously seen data, task-conditioned hypernetworks only require rehearsing task-specific weight realizations, which can be maintained in memory using a simple regularizer. Besides achieving state-of-the-art performance on standard CL benchmarks, additional experiments on long task sequences reveal that task-conditioned hypernetworks display a very large capacity to retain previous memories. Notably, such long memory lifetimes are achieved in a compressive regime, when the number of trainable hypernetwork weights is comparable or smaller than target network size. We provide insight into the structure of low-dimensional task embedding spaces (the input space of the hypernetwork) and show that task-conditioned hypernetworks demonstrate transfer learning. Finally, forward information transfer is further supported by empirical results on a challenging CL benchmark based on the CIFAR-10/100 image datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper121
- Continual learning in recurrent neural networksBenjamin Ehret, Christian Henning, Maria R. Cervera, Alexander Meulemans 等ICLR 2021 · 被引用 4,433 次
- Personalized Federated Learning using HypernetworksAviv Shamsian, Aviv Navon, Ethan Fetaya, Gal ChechikICML 2021 · 被引用 452 次
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi 等NeurIPS 2020 · 被引用 364 次
- UniControl: A Unified Diffusion Model for Controllable Visual Generation In the WildCan Qin, Shu Zhang, Ning Yu, Yihao Feng 等NeurIPS 2023 · 被引用 250 次
- HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image EditingYuval Alaluf, Omer Tov, Ron Mokady, Rinon Gal 等CVPR 2022 · 被引用 250 次
相关 Paper
- Exploiting Task Relationships in Continual Learning via Transferability-Aware Task EmbeddingsYanru Wu, Jianning Wang, Xiangyu Chen, Aurora 等NeurIPS 2025 · 被引用 3 次
- Lifelong GAN: Continual Learning for Conditional Image GenerationMengyao Zhai, Lei Chen, Frederick Tung, Jiawei He 等ICCV 2019 · 被引用 204 次
- CLR: Channel-wise Lightweight Reprogramming for Continual LearningYunhao Ge, Yuecheng Li, Shuo Ni, Jiaping Zhao 等ICCV 2023 · 被引用 16 次
- Growing a Brain with Sparsity-Inducing Generation for Continual LearningHyundong Jin, Gyeong-Hyeon Kim, Chanho Ahn, Eunwoo KimICCV 2023 · 被引用 7 次
- Is Forgetting Less a Good Inductive Bias for Forward Transfer?Jiefeng Chen, Timothy Nguyen, Dilan Görür, Arslan ChaudhryICLR 2023 · 被引用 1 次
