Lune

CVPR2021Top-tier venue

Rectification-Based Knowledge Retention for Continual Learning

Pravendra Singh, Pratik Mazumder, Piyush Rai, Vinay P. Namboodiri

2021Year
10Top-tier citations

Abstract

Deep learning models suffer from catastrophic forgetting when trained in an incremental learning setting. In this work, we propose a novel approach to address the task incremental learning problem, which involves training a model on new tasks that arrive in an incremental manner. The task incremental learning problem becomes even more challenging when the test set contains classes that are not part of the train set, i.e., a task incremental generalized zero-shot learning problem. Our approach can be used in both the zero-shot and non zero-shot task incremental learning settings. Our proposed method uses weight rectifications and affine transformations in order to adapt the model to different tasks that arrive sequentially. Specifically, we adapt the network weights to work for new tasks by "rectifying" the weights learned from the previous task. We learn these weight rectifications using very few parameters. We additionally learn affine transformations on the outputs generated by the network in order to better adapt them for the new task. We perform experiments on several datasets in both zero-shot and non zero-shot task incremental learning settings and empirically show that our approach achieves state-of-the-art results. Specifically, our approach outperforms the state-of-the-art non zero-shot task incremental learning method by over 5% on the CIFAR-100 dataset. Our approach also significantly outperforms the state-of-the-art task incremental generalized zero-shot learning method by absolute margins of 6.91% and 6.33% for the AWA1 and CUB datasets, respectively. We validate our approach using various ablation studies.

We propose a novel approach called Rectification-based Knowledge Retention (RKR) for the task incremental learning problem in the zero-shot and non zero-shot setting. Our approach (RKR) learns weight rectifications to adapt the network weights for a new task. After learning these weight rectifications, we can quickly adapt the network to work for images from that task by simply applying these weight rectifications to the network weights. We utilize an efficient technique for learning these weight rectifications to limit the model size. We also learn affine transformations (scaling factors) for all the intermediate outputs of the network that allow better adaptation of the network to the respective 1

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 5aa2e49c-b6c7-460e-86aa-c89335342f46

Cited by top-tier papers10

Ask how each one uses it

Builds on4

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines