Improving Visual Prompt Tuning for Self-supervised Vision Transformers
Seungryong Yoo, Eunji Kim, Dahuin Jung, Jungbeom Lee, Sungroh Yoon
Abstract
Visual Prompt Tuning (VPT) is an effective tuning method for adapting pretrained Vision Transformers (ViTs) to downstream tasks. It leverages extra learnable tokens, known as prompts, which steer the frozen pretrained ViTs. Although VPT has demonstrated its applicability with supervised vision transformers, it often underperforms with self-supervised ones. Through empirical observations, we deduce that the effectiveness of VPT hinges largely on the ViT blocks with which the prompt tokens interact. Specifically, VPT shows improved performance on image classification tasks for MAE and MoCo v3 when the prompt tokens are inserted into later blocks rather than the first block. These observations suggest that there exists an optimal location of blocks for the insertion of prompt tokens. Unfortunately, identifying the optimal blocks for prompts within each self-supervised ViT for diverse future scenarios is a costly process. To mitigate this problem, we propose a simple yet effective method that learns a gate for each ViT block to adjust its intervention into the prompt tokens. With our method, prompt tokens are selectively influenced by blocks that require steering for task adaptation. Our method outperforms VPT variants in FGVC and VTAB image classification and ADE20K semantic segmentation. The code is available at https://github. com/ryongithub/GatedPromptTuning .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 147bb5ef-51c8-4f9a-ae57-efb2f4238273Cited by top-tier papers25
- Visual Fourier Prompt TuningRunjia Zeng, Cheng Han, Qifan Wang, Chunshu Wu et al.NeurIPS 2024 · 58 citations
- Revisiting the Power of Prompt for Visual TuningYuzhu Wang, Lechao Cheng, Chaowei Fang, Dingwen Zhang et al.ICML 2024 · 33 citations
- Improving Visual Prompt Tuning by Gaussian Neighborhood Minimization for Long-Tailed Visual RecognitionMengke Li, Ye Liu, Yang Lu, Yiqun Zhang et al.NeurIPS 2024 · 27 citations
- Visual Instance-aware Prompt TuningXi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang et al.ACM MM 2025 · 12 citations
- FG-OrIU: Towards Better Forgetting via Feature-Gradient Orthogonality for Incremental UnlearningQian Feng, Jiahang Tu, Mintong Kang, Hanbin Zhao et al.ICCV 2025 · 9 citations
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt RelocationChikai Shang, Mengke Li, Yiqun Zhang, Zhen Chen et al.ICCV 2025 · 1 citation
- Learning Expressive Prompting With Residuals for Vision TransformersRajshekhar Das, Yonatan Dukler, Avinash Ravichandran, Ashwin SwaminathanCVPR 2023
- DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision TransformersLi Ren, Chen Chen, Liqiang Wang, Kien A. HuaCVPR 2025
- Fair-VPT: Fair Visual Prompt Tuning for Image ClassificationSungho Park, Hyeran ByunCVPR 2024 · 12 citations
- Revisit Visual Prompt Tuning: The Expressiveness of Prompt ExpertsMinh Le, Anh Nguyen, Huy Nguyen, Chau Nguyen et al.ICLR 2026 · 6 citations
