RATT: Recurrent Attention to Transient Tasks for Continual Image Captioning
Riccardo Del Chiaro, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de Weijer
Abstract
Research on continual learning has led to a variety of approaches to mitigating catastrophic forgetting in feed-forward classification networks. Until now surprisingly little attention has been focused on continual learning of recurrent models applied to problems like image captioning. In this paper we take a systematic look at continual learning of LSTM-based models for image captioning. We propose an attention-based approach that explicitly accommodates the transient nature of vocabularies in continual image captioning tasks -- i.e. that task vocabularies are not disjoint. We call our method Recurrent Attention to Transient Tasks (RATT), and also show how to adapt continual learning approaches based on weight egularization and knowledge distillation to recurrent continual learning problems. We apply our approaches to incremental image captioning problem on two new continual learning benchmarks we define using the MS-COCO and Flickr30 datasets. Our results demonstrate that RATT is able to sequentially learn five captioning tasks while incurring no forgetting of previously learned ones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e718a275-dac6-4459-ad1f-5676db870900Cited by top-tier papers20
- NICGSlowDown: Evaluating the Efficiency Robustness of Neural Image Caption Generation ModelsSimin Chen, Zihe Song, Mirazul Haque, Cong Liu et al.CVPR 2022 · 34 citations
- Decouple Before Interact: Multi-Modal Prompt Learning for Continual Visual Question AnsweringZi Qian, Xin Wang, Xuguang Duan, Pengda Qin et al.ICCV 2023 · 28 citations
- Beyond Generic: Enhancing Image Captioning with Real-World Knowledge using Vision-Language Pre-Training ModelKanzhi Cheng, Wenpo Song, Zheng Ma, Wenhao Zhu et al.ACM MM 2023 · 17 citations
- Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language TasksZijian Gao, Xingxing Zhang, Kele Xu, Xinjun Mao et al.NeurIPS 2024 · 11 citations
- Affordance-First Decomposition for Continual Learning in Video–Language UnderstandingMengzhu xu, Hanzhi Liu, Ningkang Peng, qianyu Chen et al.CVPR 2026 · 7 citations
Builds on2
Related papers
- Introducing Language Guidance in Prompt-based Continual LearningMuhammad Gul Zain Ali Khan, Muhammad Ferjad Naeem, Luc Van Gool, Didier Stricker et al.ICCV 2023 · 71 citations
- AttriCLIP: A Non-Incremental Learner for Incremental Knowledge LearningRunqi Wang, Xiaoyue Duan, Guoliang Kang, Jianzhuang Liu et al.CVPR 2023
- Knowledge Decomposition and Replay: A Novel Cross-modal Image-Text Retrieval Continual Learning MethodRui Yang, Shuang Wang, Huan Zhang, Siyuan Xu et al.ACM MM 2023 · 13 citations
- Effective Continual Learning for Text Classification with Lightweight SnapshotsJue Wang, Dajie Dong, Lidan Shou, Ke Chen et al.AAAI 2023 · 4 citations
- Continual Learning with Lifelong Vision TransformerZhen Wang, Liu Liu, Yiqun Duan, Yajing Kong et al.CVPR 2022 · 63 citations
