Boosting Visual Reprogramming for CLIP with Dual Granularity Alignment
Jiayang Wu, Xinyang Chen, Ke Lv, Weili Guan
Abstract
Model reprogramming adapts pretrained models to downstream tasks by modifying their input and output spaces. Visual reprogramming, as a prominent instance, has been explored in pioneer works on CLIP, which introduces learnable input transformations as visual prompts to repurpose its visual-language alignment for downstream visual tasks. Existing VR methods focus on single-level alignment between prompted images and text descriptions, overlooking inherent structural information in data that facilitates alignment: semantic granularity from label hierarchies and visual granularity from multi-scale representations. To address this gap, we propose Dual Granularity Alignment (DGA) with two key components for multilevel fusion. For visual granularity, we generate multiscale images and introduce Uncertainty-calibrated Prediction Fusion (UPF), which fuses predictions based on uncertainty estimation to capture hierarchical spatial information. For semantic granularity, we construct category hierarchies via Prototype-guided Label Hierarchization and develop Hierarchical Knowledge Propagation (HKP), which transfers superclass knowledge for coherent multi-level visual prompts alignment. Our DGA collaboratively integrates both granularities to enhance alignment effectiveness. Experiments across 12 downstream datasets demonstrate DGA's superiority over baselines on both ViT-based and ResNet-based CLIP architectures. Specifically, DGA achieves a 4.5% improvement over the previous state-ofthe-art method on ViT-16-based CLIP. By explicitly modeling structural granularities, DGA establishes a new paradigm for visual reprogramming. Code is available at https://github.com/JiayangWU66/DGA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f5943a5-ebfb-4ff0-849b-dd8b8a9179cfBuilds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Time-LLM: Time Series Forecasting by Reprogramming Large Language ModelsMing Jin, Shiyu Wang, Lintao Ma, Zhixuan Chu et al.ICLR 2024 · 915 citations
- Voice2Series: Reprogramming Acoustic Models for Time Series ClassificationChao-Han Huck Yang, Yun-Yun Tsai, Pin-Yu ChenICML 2021 · 150 citations
- Transfer Learning without Knowing: Reprogramming Black-box Machine Learning Models with Scarce Data and Limited ResourcesYun-Yun Tsai, Pin-Yu Chen, Tsung-Yi HoICML 2020 · 115 citations
Related papers
- Attribute-based Visual Reprogramming for Vision-Language ModelsChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi et al.ICLR 2025
- Hierarchical Cross-Modal Prompt Learning for Vision-Language ModelsHao Zheng, Shunzhi Yang, Zhuoxin He, Jinfeng Yang et al.ICCV 2025 · 5 citations
- Understanding Model Reprogramming for CLIP via Decoupling Visual PromptsChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi et al.ICML 2025
- Bayesian-guided Label Mapping for Visual ReprogrammingChengyi Cai, Zesheng Ye, Lei Feng, Jianzhong Qi et al.NeurIPS 2024 · 14 citations
- MaPLe: Multi-modal Prompt LearningMuhammad Uzair Khattak, Hanoona Abdul Rasheed, Muhammad Maaz, Salman H. Khan et al.CVPR 2023
