Time-, Memory- and Parameter-Efficient Visual Adaptation
Otniel-Bogdan Mercea, Alexey A. Gritsenko, Cordelia Schmid, Anurag Arnab
Abstract
As foundation models become more popular, there is a growing need to efficiently finetune them for downstream tasks. Although numerous adaptation methods have been proposed, they are designed to be efficient only in terms of how many parameters are trained. They, however, typi-cally still require backpropagating gradients throughout the model, meaning that their training-time and -memory cost does not reduce as significantly. We propose an adaptation method which does not back-propagate gradients through the backbone. We achieve this by designing a lightweight network in parallel that oper-ates on features from the frozen, pretrained backbone. As a result, our method is efficient not only in terms of parame-ters, but also in training-time and memory usage. Our approach achieves state-of-the-art accuracy-parameter trade-offs on the popular VTAB benchmark, and we further show how we outperform prior works with respect to training-time and -memory usage too. We further demonstrate the training efficiency and scalability of our method by adapting a vision transformer backbone of 4 billion parameters for the computationally demanding task of video classifi-cation, without any intricate model parallelism. Here, we outperform a prior adaptor-based method which could only scale to a 1 billion parameter backbone, or fully-finetuning a smaller backbone, with the same GPU and less training time. To facilitate further research, we release code at https://github.com/google-researchlscenic.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7b3072f-a4e4-4ca8-a3c2-aebe40ab0aebCited by top-tier papers14
- Frequency-Dynamic Attention Modulation for Dense PredictionLinwei Chen, Lin Gu, Ying FuICCV 2025 · 13 citations
- Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer EraFeng Lu, Tong Jin, Canming Ye, Xiangyuan Lan et al.NeurIPS 2025 · 8 citations
- MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer LearningYutong Zhang, Zimeng Wu, Shengcai Liao, Shujiang Wu et al.AAAI 2026 · 2 citations
- Svit-Split: Unleashing the Power of Vision Foundation Models Via Efficient Splitting HeadsYifan Li, Xin Li, Tianqin Li, Wenbin He et al.ICCV 2025 · 2 citations
- Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path DistillationYutong Zhang, Jiaxin Chen, Honglin Chen, Kaiqi Zheng et al.CVPR 2026 · 1 citation
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- AIM: Adapting Image Models for Efficient Video Action RecognitionTaojiannan Yang, Yi Zhu, Yusheng Xie, Aston Zhang et al.ICLR 2023 · 62 citations
- Parameter-efficient is not Sufficient: Exploring Parameter, Memory, and Time Efficient Adapter Tuning for Dense PredictionsDongshuo Yin, Xueting Han, Bin Li, Hao Feng et al.ACM MM 2024 · 18 citations
- Parameter-, Memory-, Time-Efficient Multi-Task Dense Vision AdaptationHaiming Yao, Wei Luo, Qiyu Chen, Jianxing Liao et al.AAAI 2026
- Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation ModelsBenjamin Ramtoula, Pierre-Yves Lajoie, Paul Newman, Daniele De MartiniNeurIPS 2025 · 2 citations
- Visual Query Tuning: Towards Effective Usage of Intermediate Representations for Parameter and Memory Efficient Transfer LearningCheng-Hao Tu, Zheda Mai, Wei-Lun ChaoCVPR 2023
