Learning Expressive Prompting With Residuals for Vision Transformers
Rajshekhar Das, Yonatan Dukler, Avinash Ravichandran, Ashwin Swaminathan
Abstract
Prompt learning is an efficient approach to adapt transformers by inserting learnable set of parameters into the input and intermediate representations of a pre-trained model. In this work, we present Expressive Prompts with Residuals (EXPRES) which modifies the prompt learning paradigm specifically for effective adaptation of vision transformers (ViT). Our method constructs downstream representations via learnable "output" tokens (shallow prompts), that are akin to the learned class tokens of the ViT. Further for better steering of the downstream representation processed by the frozen transformer, we introduce residual learnable tokens that are added to the output of various computations. We apply EXPRES for image classification and few-shot semantic segmentation, and show our method is capable of achieving state of the art prompt tuning on 3/3 categories of the VTAB benchmark. In addition to strong performance, we observe that our approach is an order of magnitude more prompt efficient than existing visual prompting baselines. We analytically show the computational benefits of our approach over weight space adaptation techniques like finetuning. Lastly we systematically corroborate the architectural design of our method via a series of ablation experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c4a162d3-5f81-4185-be75-6f49a3232e16Cited by top-tier papers13
- Visual Fourier Prompt TuningRunjia Zeng, Cheng Han, Qifan Wang, Chunshu Wu et al.NeurIPS 2024 · 58 citations
- SA²VP: Spatially Aligned-and-Adapted Visual PromptWenjie Pei, Tongqi Xia, Fanglin Chen, Jinsong Li et al.AAAI 2024 · 33 citations
- Visual Instance-aware Prompt TuningXi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang et al.ACM MM 2025 · 12 citations
- Compressed Video Prompt TuningBing Li, Jiaxin Chen, Xiuguo Bao, Di HuangNeurIPS 2023 · 11 citations
- Time-, Memory- and Parameter-Efficient Visual AdaptationOtniel-Bogdan Mercea, Alexey A. Gritsenko, Cordelia Schmid, Anurag ArnabCVPR 2024 · 11 citations
Builds on40
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- PANet: Few-Shot Image Semantic Segmentation With Prototype AlignmentKaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou et al.ICCV 2019 · 1,404 citations
Related papers
- Improving Visual Prompt Tuning for Self-supervised Vision TransformersSeungryong Yoo, Eunji Kim, Dahuin Jung, Jungbeom Lee et al.ICML 2023 · 74 citations
- Fair-VPT: Fair Visual Prompt Tuning for Image ClassificationSungho Park, Hyeran ByunCVPR 2024 · 12 citations
- DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision TransformersLi Ren, Chen Chen, Liqiang Wang, Kien A. HuaCVPR 2025
- E2VPT: An Effective and Efficient Approach for Visual Prompt TuningCheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao et al.ICCV 2023 · 108 citations
- Revisit Visual Prompt Tuning: The Expressiveness of Prompt ExpertsMinh Le, Anh Nguyen, Huy Nguyen, Chau Nguyen et al.ICLR 2026 · 6 citations
