Token Coordinated Prompt Attention is Needed for Visual Prompting
Zichen Liu, Xu Zou, Gang Hua, Jiahuan Zhou
摘要
Visual prompting techniques are widely used to efficiently fine-tune pretrained Vision Transformers (ViT) by learning a small set of shared prompts for all tokens. However, existing methods overlook the unique roles of different tokens in conveying discriminative information and interact with all tokens using the same prompts, thereby limiting the representational capacity of ViT. This often leads to indistinguishable and biased prompt-extracted features, hindering performance. To address this issue, we propose a plug-and-play Token Coordinated Prompt Attention (TCPA) module, which assigns specific coordinated prompts to different tokens for attention-based interactions. Firstly, recognizing the distinct functions of CLS and image tokens-global information aggregation and local feature extraction, we disentangle the prompts into CLS Prompts and Image Prompts, which interact exclusively with CLS tokens and image tokens through attention mechanisms. This enhances their respective discriminative abilities. Furthermore, as different image tokens correspond to distinct image patches and contain diverse information, we employ a matching function to automatically assign coordinated prompts to individual tokens. This enables more precise attention interactions, improving the diversity and representational capacity of the extracted features. Extensive experiments across various benchmarks demonstrate that TCPA significantly enhances the diversity and discriminative power of the extracted features. The code is available at https://github.com/zhoujiahuan1991/ICML2025-TCPA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- VIPAMIN: Visual Prompt Initialization via Embedding Selection and Subspace ExpansionJaekyun Park, Hye Won ChungNeurIPS 2025 · 被引用 1 次
- Parameters as Experts: Adapting Vision Models with Dynamic Parameter RoutingMeng Lou, Stanley Yu, Yizhou YuICML 2026 · 被引用 1 次
- Boosting Visual Reprogramming for CLIP with Dual Granularity AlignmentJiayang Wu, Xinyang Chen, Ke Lv, Weili GuanCVPR 2026
- UniExtreme: A Universal Foundation Model for Extreme Weather ForecastingHang Ni, Weijia Zhang, Hao LiuKDD 2026
它引用的顶会 Paper21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
相关 Paper
- Fair-VPT: Fair Visual Prompt Tuning for Image ClassificationSungho Park, Hyeran ByunCVPR 2024 · 被引用 12 次
- DA-VPT: Semantic-Guided Visual Prompt Tuning for Vision TransformersLi Ren, Chen Chen, Liqiang Wang, Kien A. HuaCVPR 2025
- Improving Visual Prompt Tuning for Self-supervised Vision TransformersSeungryong Yoo, Eunji Kim, Dahuin Jung, Jungbeom Lee 等ICML 2023 · 被引用 74 次
- Prompt-CAM: Making Vision Transformers Interpretable for Fine-Grained AnalysisArpita Chowdhury, Dipanjyoti Paul, Zheda Mai, Jianyang Gu 等CVPR 2025
- Selective Visual Prompting in Vision MambaYifeng Yao, Zichen Liu, Zhenyu Cui, Yuxin Peng 等AAAI 2025 · 被引用 13 次
