Attribute-based Visual Reprogramming for Vision-Language Models
Chengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi, Feng Liu
Abstract
Visual reprogramming (VR) reuses pre-trained vision models for downstream image classification tasks by adding trainable noise patterns to inputs. When applied to vision-language models (e.g., CLIP), existing VR approaches follow the same pipeline used in vision models (e.g., ResNet, ViT), where ground-truth class labels are inserted into fixed text templates to guide the optimization of VR patterns. This label-based approach, however, overlooks the rich information and diverse attribute-guided textual representations that CLIP can exploit, which may lead to the misclassification of samples. In this paper, we propose Attribute-based Visual Reprogramming (AttrVR) for CLIP, utilizing descriptive attributes (DesAttrs) and distinctive attributes (DistAttrs), which respectively represent common and unique feature descriptions for different classes. Besides, as images of the same class may reflect different attributes after VR, AttrVR iteratively refines patterns using the -nearest DesAttrs and DistAttrs for each image sample, enabling more dynamic and sample-specific optimization. Theoretically, AttrVR is shown to reduce intra-class variance and increase inter-class separation. Empirically, it achieves superior performance in 12 downstream tasks for both ViT-based and ResNet-based CLIP. The success of AttrVR facilitates more effective integration of VR from unimodal vision models into vision-language models. Our code is available at https://github.com/tmlr-group/AttrVR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ecb0be68-86fd-4678-b75c-b748b0911622Cited by top-tier papers3
- LOREAL: Mitigating Low-Resolution Challenges in Vision-Language Models with Attribute-driven Prompt Self-DistillationXucong Wang, Pengkun Wang, Zhe Zhao, Liheng Yu et al.CVPR 2026
- Leveraging Evidence Priors for Robust Prompt Learning under Noisy Supervision in Vision-Language ModelsJunnan Zou, Zhu Teng, Wei Zhang, Ming He et al.ICML 2026
- Boosting Visual Reprogramming for CLIP with Dual Granularity AlignmentJiayang Wu, Xinyang Chen, Ke Lv, Weili GuanCVPR 2026
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- What does a platypus look like? Generating customized prompts for zero-shot image classificationSarah M. Pratt, Ian Covert, Rosanne Liu, Ali FarhadiICCV 2023 · 343 citations
- Voice2Series: Reprogramming Acoustic Models for Time Series ClassificationChao-Han Huck Yang, Yun-Yun Tsai, Pin-Yu ChenICML 2021 · 150 citations
- Transfer Learning without Knowing: Reprogramming Black-box Machine Learning Models with Scarce Data and Limited ResourcesYun-Yun Tsai, Pin-Yu Chen, Tsung-Yi HoICML 2020 · 115 citations
Related papers
- Understanding Model Reprogramming for CLIP via Decoupling Visual PromptsChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi et al.ICML 2025
- Bayesian-guided Label Mapping for Visual ReprogrammingChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi et al.NeurIPS 2024 · 14 citations
- Understanding and Improving Visual Prompting: A Label-Mapping PerspectiveAochuan Chen, Yuguang Yao, Pin-Yu Chen, Yihua Zhang et al.CVPR 2023
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- Advancing Textual Prompt Learning with Anchored AttributesZheng Li, Yibing Song, Ming-Ming Cheng, Xiang Li et al.ICCV 2025 · 8 citations
