InsVP: Efficient Instance Visual Prompting from Image Itself
Zichen Liu, Yuxin Peng, Jiahuan Zhou
摘要
Visual prompting is an efficient methodology for finetuning pretrained visual models by introducing a small number of learnable parameters while keeping the backbone frozen. However, most existing visual prompting methods learn a shared prompt for all samples, making it challenging to grasp distinct characteristics among diverse samples, thereby limiting the model's performance. While other methods partially address this issue through sample clustering and learning multiple prompts, they still struggle to capture nuanced differences among instances and incur significant parameter overhead. Therefore, to comprehensively and efficiently leverage discriminative characteristics of individual instances, we propose an Instance Visual Prompting method, called InsVP. Initially, the instance image prompt is introduced to extract both crucial and nuanced discriminative information from the original image itself and is overlaid onto the input image. Furthermore, the instance feature prompt is designed to capture both commonalities and characteristics among individual instances, fed into the model's intermediate layers to facilitate feature extraction. Consequently, the instance image and feature prompts complement each other, enhancing the adaptation ability of pretrained models to extract discriminative features from individual instances. Extensive experiments on various large-scale benchmarks show that our InsVP achieves superior performance exceeding the state-ofthe-art methods at a lower parameter cost. The code is available at https://github.com/zhoujiahuan1991/MM2024-InsVP
• Computing methodologies → Computer vision.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Visual Instance-aware Prompt TuningXi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang 等ACM MM 2025 · 被引用 12 次
- State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video UnderstandingJiahuan Zhou, Kai Zhu, Zhenyu Cui, Zichen Liu 等NeurIPS 2025 · 被引用 2 次
- Revisiting Pool-Based Prompt Learning for Few-Shot Class-Incremental LearningYongwei Jiang, Yixiong Zou, Yuhua Li, Ruixuan LiICCV 2025 · 被引用 1 次
- UPP: Unified Point-Level Prompting for Robust Point Cloud AnalysisZixiang Ai, Zhenyu Cui, Yuxin Peng, Jiahuan ZhouICCV 2025 · 被引用 1 次
- Componential Prompt-Knowledge Alignment for Domain Incremental LearningKunlun Xu, Xu Zou, Gang Hua, Jiahuan ZhouICML 2025
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- Selective Visual Prompting in Vision MambaYifeng Yao, Zichen Liu, Zhenyu Cui, Yuxin Peng 等AAAI 2025 · 被引用 13 次
- Diversity-Aware Meta Visual PromptingQidong Huang, Xiaoyi Dong, Dongdong Chen, Weiming Zhang 等CVPR 2023
- LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model AdaptationCan Jin, Ying Li, Mingyu Zhao, Shiyu Zhao 等ICLR 2025
- SA²VP: Spatially Aligned-and-Adapted Visual PromptWenjie Pei, Tongqi Xia, Fanglin Chen, Jinsong Li 等AAAI 2024 · 被引用 33 次
- AutoVP: An Automated Visual Prompting Framework and BenchmarkHsi-Ai Tsao, Lei Hsiung, Pin-Yu Chen, Si Liu 等ICLR 2024 · 被引用 29 次
