Lune

CVPR2026顶会

Boosting Visual Reprogramming for CLIP with Dual Granularity Alignment

Jiayang Wu, Xinyang Chen, Ke Lv, Weili Guan

出版方
2026年份

摘要

Model reprogramming adapts pretrained models to downstream tasks by modifying their input and output spaces. Visual reprogramming, as a prominent instance, has been explored in pioneer works on CLIP, which introduces learnable input transformations as visual prompts to repurpose its visual-language alignment for downstream visual tasks. Existing VR methods focus on single-level alignment between prompted images and text descriptions, overlooking inherent structural information in data that facilitates alignment: semantic granularity from label hierarchies and visual granularity from multi-scale representations. To address this gap, we propose Dual Granularity Alignment (DGA) with two key components for multilevel fusion. For visual granularity, we generate multiscale images and introduce Uncertainty-calibrated Prediction Fusion (UPF), which fuses predictions based on uncertainty estimation to capture hierarchical spatial information. For semantic granularity, we construct category hierarchies via Prototype-guided Label Hierarchization and develop Hierarchical Knowledge Propagation (HKP), which transfers superclass knowledge for coherent multi-level visual prompts alignment. Our DGA collaboratively integrates both granularities to enhance alignment effectiveness. Experiments across 12 downstream datasets demonstrate DGA's superiority over baselines on both ViT-based and ResNet-based CLIP architectures. Specifically, DGA achieves a 4.5% improvement over the previous state-ofthe-art method on ViT-16-based CLIP. By explicitly modeling structural granularities, DGA establishes a new paradigm for visual reprogramming. Code is available at https://github.com/JiayangWU66/DGA.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖