Selective Vision-Language Subspace Projection for Few-shot CLIP
Xingyu Zhu, Beier Zhu, Yi Tan, Shuo Wang, Yanbin Hao, Hanwang Zhang
Abstract
Vision-language models such as CLIP are capable of mapping the different modality data into a unified feature space, enabling zero/fewshot inference by measuring the similarity of given images and texts. However, most existing methods overlook modality gaps in CLIP's encoded features, which is shown as the text and image features lie far apart from each other, resulting in limited classification performance. To tackle this issue, we introduce a method called Selective Vision-Language Subspace Projection (SSP), which incorporates local image features and utilizes them as a bridge to enhance the alignment between image-text pairs. Specifically, our SSP framework comprises two parallel modules: a vision projector and a language projector. Both projectors utilize local image features to span the respective subspaces for image and texts, thereby projecting the image and text features into their respective subspaces to achieve alignment. Moreover, our approach entails only training-free matrix calculations and can be seamlessly integrated into advanced CLIPbased few-shot learning frameworks. Extensive experiments on 11 datasets have demonstrated SSP's superior text-image alignment capabilities, outperforming the state-of-the-art alignment methods. The code is available at https://github.com/zhuhsingyuu/SSP
• Computing methodologies → Computer vision representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c177fca8-e72e-4268-bfb4-45cea2f79c09Cited by top-tier papers19
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward RectificationYongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng et al.ICLR 2026 · 130 citations
- Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias CorrectingXingyu Zhu, Beier Zhu, Yi Tan, Shuo Wang et al.NeurIPS 2024 · 36 citations
- Enhancing CLIP Robustness via Cross-Modality AlignmentXingyu Zhu, Beier Zhu, Shuo Wang, Kesen Zhao et al.NeurIPS 2025 · 17 citations
- Real-Time Motion-Controllable Autoregressive Video DiffusionKesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou et al.ICLR 2026 · 10 citations
- Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination MitigationXingyu Zhu, Kesen Zhao, Liang Yi, Shuo Wang et al.ICLR 2026 · 9 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language ModelsJinhao Li, Haopeng Li, Sarah Monazam Erfani, Lei Feng et al.ICML 2024 · 30 citations
- FewVS: A Vision-Semantics Integration Framework for Few-Shot Image ClassificationZhuoling Li, Yong Wang, Kaitong LiACM MM 2024 · 4 citations
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang et al.ICML 2024 · 7 citations
- IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal AlignmentSimone Magistri, Dipam Goswami, Marco Mistretta, Bartlomiej Twardowski et al.CVPR 2026 · 4 citations
- Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language ModelShiming Chen, Bowen Duan, Salman Khan, Fahad Shahbaz KhanICCV 2025 · 2 citations
