Not All Features Matter: Enhancing Few-shot CLIP with Adaptive Prior Refinement
Xiangyang Zhu, Renrui Zhang, Bowei He, Aojun Zhou, Dong Wang, Bin Zhao, Peng Gao
Abstract
The popularity of Contrastive Language-Image Pretraining (CLIP) has propelled its application to diverse downstream vision tasks. To improve its capacity on downstream tasks, few-shot learning has become a widely-adopted technique. However, existing methods either exhibit limited performance or suffer from excessive learnable parameters. In this paper, we propose APE, an Adaptive Prior rEfinement method for CLIP's pre-trained knowledge, which achieves superior accuracy with high computational efficiency. Via a prior refinement module, we analyze the inter-class disparity in the downstream data and decouple the domain-specific knowledge from the CLIP-extracted cache model. On top of that, we introduce two model variants, a training-free APE and a training-required APE-T. We explore the trilateral affinities between the test image, prior cache model, and textual representations, and only enable a lightweight category-residual module to be trained. For the average accuracy over 11 benchmarks, both APE and APE-T attain state-of-the-art and respectively outperform the second-best by +1.59% and +1.99% under 16 shots with ×30 less learnable parameters. Code is available at https://github.com/yangyangyang127/APE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers51
- GraphAdapter: Tuning Vision-Language Models With Dual Knowledge GraphXin Li, Dongze Lian, Zhihe Lu, Jiawang Bai et al.NeurIPS 2023 · 138 citations
- Referred by Multi-Modality: A Unified Temporal Transformer for Video Object SegmentationShilin Yan, Renrui Zhang, Ziyu Guo, Wenchao Chen et al.AAAI 2024 · 67 citations
- Meta-Adapter: An Online Few-shot Learner for Vision-Language ModelCheng Cheng, Lin Song, Ruoyi Xue, Hang Wang et al.NeurIPS 2023 · 65 citations
- A Hard-to-Beat Baseline for Training-free CLIP-based AdaptationZhengbo Wang, Jian Liang, Lijun Sheng, Ran He et al.ICLR 2024 · 51 citations
- AWT: Transferring Vision-Language Models via Augmentation, Weighting, and TransportationYuhan Zhu, Yuyang Ji, Zhiyu Zhao, Gangshan Wu et al.NeurIPS 2024 · 45 citations
Builds on33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free AttentionZiyu Guo, Renrui Zhang, Longtian Qiu, Xianzheng Ma et al.AAAI 2023 · 182 citations
- Rethinking Prior Information Generation with CLIP for Few-Shot SegmentationJin Wang, Bingfeng Zhang, Jian Pang, Honglong Chen et al.CVPR 2024 · 27 citations
- Seeing in Flowing: Adapting CLIP for Action Recognition with Motion Prompts LearningQiang Wang, Junlong Du, Ke Yan, Shouhong DingACM MM 2023 · 26 citations
- RA-CLIP: Retrieval Augmented Contrastive Language-Image Pre-TrainingChen-Wei Xie, Siyang Sun, Xiong Xiong, Yun Zheng et al.CVPR 2023
- Decoupling Template Bias in CLIP: Harnessing Empty Prompts for Enhanced Few-Shot LearningZhenyu Zhang, Guangyao Chen, Yixiong Zou, Zhimeng Huang et al.AAAI 2026 · 3 citations
