DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision Transformers
Li Ren, Chen Chen, Liqiang Wang, Kien A. Hua
摘要
Visual Prompt Tuning (VPT) has become a promising solution for Parameter-Efficient Fine-Tuning (PEFT) approach for Vision Transformer (ViT) models by partially fine-tuning learnable tokens while keeping most model parameters frozen. Recent research has explored modifying the connection structures of the prompts. However, the fundamental correlation and distribution between the prompts and image tokens remain unexplored. In this paper, we leverage metric learning techniques to investigate how the distribution of prompts affects fine-tuning performance. Specifically, we propose a novel framework, Distribution Aware Visual Prompt Tuning (DA-VPT), to guide the distributions of the prompts by learning the distance metric from their class-related semantic data. Our method demonstrates that the prompts can serve as an effective bridge to share semantic information between image patches and the class token. We extensively evaluated our approach on popular benchmarks in both recognition and segmentation tasks. The results demonstrate that our approach enables more effective and efficient fine-tuning of ViT models by leveraging semantic information to guide the learning of the prompts, leading to improved performance on various downstream vision tasks. The code is released on https://github.com/Noahsark/DA-VPT . Tunable Frozen Image Token Prompt Pull Push CLS Token Transformer Block Layer CLS HEAD … 𝐿 𝐶𝐸 … CLS Transformer Block Layer CLS … … VPT-Deep Transformer Block Layer CLS HEAD … 𝐿 𝐶𝐸 … CLS Transformer Block Layer CLS … … 𝐿 𝑀𝐿 ( , ) 𝐿 𝑀𝐿 ( , )
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Visual Instance-aware Prompt TuningXi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang 等ACM MM 2025 · 被引用 12 次
- Visual Prompt-Agnostic EvolutionJunze Wang, Lei Fan, Dezheng Zhang, Weipeng Jing 等ICLR 2026 · 被引用 4 次
- VIPAMIN: Visual Prompt Initialization via Embedding Selection and Subspace ExpansionJaekyun Park, Hye Won ChungNeurIPS 2025 · 被引用 1 次
- SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View FusionXiang Li, Heqian Qiu, Lanxiao Wang, Benliu Qiu 等CVPR 2026 · 被引用 1 次
- LoPrune: Efficient Data Pruning for LoRA-Based Fine-Tuning of Vision TransformerQiang He, Yaozong Yang, KAIBIN WANG, Ziteng Wei 等CVPR 2026
它引用的顶会 Paper30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
相关 Paper
- Learning Semantic Proxies from Visual Prompts for Parameter-Efficient Fine-Tuning in Deep Metric LearningLi Ren, Chen Chen, Liqiang Wang, Kien A. HuaICLR 2024 · 被引用 7 次
- Improving Visual Prompt Tuning for Self-supervised Vision TransformersSeungryong Yoo, Eunji Kim, Dahuin Jung, Jungbeom Lee 等ICML 2023 · 被引用 74 次
- Fair-VPT: Fair Visual Prompt Tuning for Image ClassificationSungho Park, Hyeran ByunCVPR 2024 · 被引用 12 次
- CVPT: Cross Visual Prompt TuningLingyun Huang, Jianxu Mao, Junfei Yi, Ziming Tao 等ICCV 2025 · 被引用 6 次
- PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt RelocationChikai Shang, Mengke Li, Yiqun Zhang, Zhen Chen 等ICCV 2025 · 被引用 1 次
