Sample-specific Masks for Visual Reprogramming-based Prompting
Chengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi, Feng Liu
Abstract
Visual reprogramming (VR) is a prompting technique that aims to re-purpose a pre-trained model (e.g., a classifier on ImageNet) to target tasks (e.g., medical data prediction) by learning a small-scale pattern added into input images instead of tuning considerable parameters within the model. The location of the pattern within input samples is usually determined by a pre-defined mask shared across all samples. In this paper, we show that the shared mask potentially limits VR's generalization and increases its approximation error due to the lack of sample-level adaptation. Motivated by this finding, we design a new framework for VR called sample-specific multi-channel masks (SMM). Specifically, SMM employs a lightweight ConvNet and patch-wise interpolation to generate sample-specific three-channel masks instead of a shared and pre-defined mask. Since we generate different masks for individual samples, SMM is theoretically shown to reduce approximation error for the target tasks compared with existing state-of-the-art VR methods. We also empirically demonstrate its performance gain on both ResNet and ViT. The success of SMM further highlights the broader applicability of VR in leveraging the latent knowledge of pre-trained models for various target tasks. Our code is available at https: //github.com/tmlr-group/SMM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake DetectionKaiqing Lin, Yuzhen Lin, Weixiang Li, Taiping Yao et al.AAAI 2025 · 32 citations
- Bayesian-guided Label Mapping for Visual ReprogrammingChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi et al.NeurIPS 2024 · 14 citations
- Prime Once, then Reprogram Locally: An Efficient Alternative to Black-Box Service Model AdaptationYunbei Zhang, Chengyi Cai, Feng Liu, Jihun HammCVPR 2026 · 5 citations
- Generalization-Preserved Learning: Closing the Backdoor to Catastrophic Forgetting in Continual Deepfake DetectionXueyi Zhang, Peiyin Zhu, Chengwei Zhang, Zhiyuan Yan et al.ICCV 2025 · 3 citations
- Enhancing Visual Prompting through Expanded Transformation Space and Overfitting MitigationShohei EnomotoNeurIPS 2025 · 1 citation
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Voice2Series: Reprogramming Acoustic Models for Time Series ClassificationChao-Han Huck Yang, Yun-Yun Tsai, Pin-Yu ChenICML 2021 · 150 citations
Related papers
- Attribute-based Visual Reprogramming for Vision-Language ModelsChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi et al.ICLR 2025
- Understanding Model Reprogramming for CLIP via Decoupling Visual PromptsChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi et al.ICML 2025
- Endowing Visual Reprogramming with Adversarial RobustnessShengjie Zhou, Xin Cheng, Haiyang Xu, Ming Yan et al.ICLR 2025
- MEDICAL IMAGE UNDERSTANDING WITH PRETRAINED VISION LANGUAGE MODELS: A COMPREHENSIVE STUDYZiyuan Qin, Huahui Yi, Qicheng Lao, Kang LiICLR 2023 · 25 citations
- Visual Prompting via Image InpaintingAmir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson et al.NeurIPS 2022 · 340 citations
