Fine-tuning Image Transformers using Learnable Memory
Mark Sandler, Andrey Zhmoginov, Max Vladymyrov, Andrew Jackson
Abstract
In this paper we propose augmenting Vision Transformer models with learnable memory tokens. Our approach allows the model to adapt to new tasks, using few parameters, while optionally preserving its capabilities on previously learned tasks. At each layer we introduce a set of learnable embedding vectors that provide contextual information useful for specific datasets. We call these “memory tokens”. We show that augmenting a model with just a handful of such tokens per layer significantly improves accuracy when compared to conventional head-only fine-tuning, and performs only slightly below the significantly more expensive full fine-tuning. We then propose an attention-masking approach that enables extension to new downstream tasks, with a computation reuse. In this setup in addition to being parameters efficient, models can execute both old and new tasks as a part of single inference at a small incremental cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9ec0f7b-cd58-4e05-91f6-d5c9c2ecf635Cited by top-tier papers17
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 769 citations
- Bayesian Prompt Learning for Image-Language Model GeneralizationMohammad Mahdi Derakhshani, Enrique Sanchez, Adrian Bulat, Victor Guilherme Turrisi da Costa et al.ICCV 2023 · 66 citations
- MasQCLIP for Open-Vocabulary Universal Image SegmentationXin Xu, Tianyi Xiong, Zheng Ding, Zhuowen TuICCV 2023 · 57 citations
- Fine-tuning Reinforcement Learning Models is Secretly a Forgetting Mitigation ProblemMaciej Wolczyk, Bartlomiej Cupial, Mateusz Ostaszewski, Michal Bortkiewicz et al.ICML 2024 · 29 citations
- Spectral Prompt Tuning: Unveiling Unseen Classes for Zero-Shot Semantic SegmentationWenhao Xu, Rongtao Xu, Changwei Wang, Shibiao Xu et al.AAAI 2024 · 21 citations
Builds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
Related papers
- Visual Query Tuning: Towards Effective Usage of Intermediate Representations for Parameter and Memory Efficient Transfer LearningCheng-Hao Tu, Zheda Mai, Wei-Lun ChaoCVPR 2023
- Task Adaptive Parameter Sharing for Multi-Task LearningMatthew Wallingford, Hao Li, Alessandro Achille, Avinash Ravichandran et al.CVPR 2022 · 61 citations
- Token Mixing: Parameter-Efficient Transfer Learning from Image-Language to Video-LanguageYuqi Liu, Luhui Xu, Pengfei Xiong, Qin JinAAAI 2023 · 10 citations
- Dynamic Tuning Towards Parameter and Inference Efficiency for ViT AdaptationWangbo Zhao, Jiasheng Tang, Yizeng Han, Yibing Song et al.NeurIPS 2024 · 41 citations
- Memory Layers at ScaleVincent-Pierre Berges, Barlas Oguz, Daniel Haziza, Wen-tau Yih et al.ICML 2025
