Sample-specific Masks for Visual Reprogramming-based Prompting
Chengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi, Feng Liu
摘要
Visual reprogramming (VR) is a prompting technique that aims to re-purpose a pre-trained model (e.g., a classifier on ImageNet) to target tasks (e.g., medical data prediction) by learning a small-scale pattern added into input images instead of tuning considerable parameters within the model. The location of the pattern within input samples is usually determined by a pre-defined mask shared across all samples. In this paper, we show that the shared mask potentially limits VR's generalization and increases its approximation error due to the lack of sample-level adaptation. Motivated by this finding, we design a new framework for VR called sample-specific multi-channel masks (SMM). Specifically, SMM employs a lightweight ConvNet and patch-wise interpolation to generate sample-specific three-channel masks instead of a shared and pre-defined mask. Since we generate different masks for individual samples, SMM is theoretically shown to reduce approximation error for the target tasks compared with existing state-of-the-art VR methods. We also empirically demonstrate its performance gain on both ResNet and ViT. The success of SMM further highlights the broader applicability of VR in leveraging the latent knowledge of pre-trained models for various target tasks. Our code is available at https: //github.com/tmlr-group/SMM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Standing on the Shoulders of Giants: Reprogramming Visual-Language Model for General Deepfake DetectionKaiqing Lin, Yuzhen Lin, Weixiang Li, Taiping Yao 等AAAI 2025 · 被引用 32 次
- Bayesian-guided Label Mapping for Visual ReprogrammingChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi 等NeurIPS 2024 · 被引用 14 次
- Prime Once, then Reprogram Locally: An Efficient Alternative to Black-Box Service Model AdaptationYunbei Zhang, Chengyi Cai, Feng Liu, Jihun HammCVPR 2026 · 被引用 5 次
- Generalization-Preserved Learning: Closing the Backdoor to Catastrophic Forgetting in Continual Deepfake DetectionXueyi Zhang, Peiyin Zhu, Chengwei Zhang, Zhiyuan Yan 等ICCV 2025 · 被引用 3 次
- Enhancing Visual Prompting through Expanded Transformation Space and Overfitting MitigationShohei EnomotoNeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- Voice2Series: Reprogramming Acoustic Models for Time Series ClassificationChao-Han Huck Yang, Yun-Yun Tsai, Pin-Yu ChenICML 2021 · 被引用 150 次
相关 Paper
- Attribute-based Visual Reprogramming for Vision-Language ModelsChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi 等ICLR 2025
- Understanding Model Reprogramming for CLIP via Decoupling Visual PromptsChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi 等ICML 2025
- Endowing Visual Reprogramming with Adversarial RobustnessShengjie Zhou, Xin Cheng, Haiyang Xu, Ming Yan 等ICLR 2025
- MEDICAL IMAGE UNDERSTANDING WITH PRETRAINED VISION LANGUAGE MODELS: A COMPREHENSIVE STUDYZiyuan Qin, Huahui Yi, Qicheng Lao, Kang LiICLR 2023 · 被引用 25 次
- Visual Prompting via Image InpaintingAmir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson 等NeurIPS 2022 · 被引用 340 次
