Facing the Elephant in the Room: Visual Prompt Tuning or Full finetuning?
Cheng Han, Qifan Wang, Yiming Cui, Wenguan Wang, Lifu Huang, Siyuan Qi, Dongfang Liu
摘要
As the scale of vision models continues to grow, the emergence of Visual Prompt Tuning (VPT) as a parameter-efficient transfer learning technique has gained attention due to its superior performance compared to traditional full-finetuning. However, the conditions favoring VPT (the when") and the underlying rationale (the why") remain unclear. In this paper, we conduct a comprehensive analysis across 19 distinct datasets and tasks. To understand the when"aspect, we identify the scenarios where VPT proves favorable by two dimensions: task objectives and data distributions. We find that VPT is preferrable when there is 1) a substantial disparity between the original and the downstream task objectives (e.g., transitioning from classification to counting), or 2) a similarity in data distributions between the two tasks (e.g., both involve natural images). In exploring the why"dimension, our results indicate VPT's success cannot be attributed solely to overfitting and optimization considerations. The unique way VPT preserves original features and adds parameters appears to be a pivotal factor. Our study provides insights into VPT's mechanisms, and offers guidance for its optimal utilization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Visual Fourier Prompt TuningRunjia Zeng, Cheng Han, Qifan Wang, Chunshu Wu 等NeurIPS 2024 · 被引用 58 次
- M²PT: Multimodal Prompt Tuning for Zero-shot Instruction LearningTaowen Wang, Yiyang Liu, James Liang, Junhan Zhao 等EMNLP 2024 · 被引用 31 次
- Improving Visual Prompt Tuning by Gaussian Neighborhood Minimization for Long-Tailed Visual RecognitionMengke Li, Ye Liu, Yang Lu, Yiqun Zhang 等NeurIPS 2024 · 被引用 27 次
- All You Need is One: Capsule Prompt Tuning with a Single VectorYiyang Liu, James Liang, Heng Fan, Wenhao Yang 等NeurIPS 2025 · 被引用 16 次
- Visual Instance-aware Prompt TuningXi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang 等ACM MM 2025 · 被引用 12 次
它引用的顶会 Paper36
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision TransformersLi Ren, Chen Chen, Liqiang Wang, Kien A. HuaCVPR 2025
- π-Tuning: Transferring Multimodal Foundation Models with Optimal Multi-task InterpolationChengyue Wu, Teng Wang, Yixiao Ge, Zeyu Lu 等ICML 2023 · 被引用 50 次
- E2VPT: An Effective and Efficient Approach for Visual Prompt TuningCheng Han, Qifan Wang, Yiming Cui, Zhiwen Cao 等ICCV 2023 · 被引用 108 次
- Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters ThemselvesShihan Wu, Ji Zhang, Pengpeng Zeng, Lianli Gao 等CVPR 2025
- VMT-Adapter: Parameter-Efficient Transfer Learning for Multi-Task Dense Scene UnderstandingYi Xin, Junlong Du, Qiang Wang, Zhiwen Lin 等AAAI 2024 · 被引用 94 次
