Post-pre-training for Modality Alignment in Vision-Language Foundation Models
Shin'ya Yamaguchi, Dewei Feng, Sekitoshi Kanai, Kazuki Adachi, Daiki Chijiwa
Abstract
Contrastive language image pre-training (CLIP) is an essential component of building modern vision-language foundation models. While CLIP demonstrates remarkable zeroshot performance on downstream tasks, the multi-modal feature spaces still suffer from a modality gap, which is a gap between image and text feature clusters and limits downstream task performance. Although existing works attempt to address the modality gap by modifying pre-training or fine-tuning, they struggle with heavy training costs with large datasets or degradations of zero-shot performance. This paper presents CLIP-Refine, a post-pre-training method for CLIP models at a phase between pre-training and finetuning. CLIP-Refine aims to align the feature space with 1 epoch training on small image-text datasets without zeroshot performance degradations. To this end, we introduce two techniques: random feature alignment (RaFA) and hybrid contrastive-distillation (HyCD). RaFA aligns the image and text features to follow a shared prior distribution by minimizing the distance to random reference vectors sampled from the prior. HyCD updates the model with hybrid soft labels generated by combining ground-truth image-text pair labels and outputs from the pre-trained CLIP model. This contributes to achieving both maintaining the past knowledge and learning new knowledge to align features. Our extensive experiments with multiple classification and retrieval tasks show that CLIP-Refine succeeds in mitigating the modality gap and improving the zero-shot performance 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6663ff62-6762-4591-a397-a0760f8586a2Cited by top-tier papers6
- T-REGS: Minimum Spanning Tree Regularization for Self-Supervised LearningJulie Mordacq, David Loiseaux, Vicky Kalogeiton, Steve OudotNeurIPS 2025 · 2 citations
- CHIPS: Efficient CLIP Adaptation via Curvature-aware Hybrid Influence-based Data SelectionXinlin Zhuang, Yichen Li, Xiwei Liu, Haolin Yang et al.CVPR 2026 · 1 citation
- Piercing the Fog: Disentangling Key Features for Vision Models in Multi-Degradation ScenariosSiyu Chen, Shiqiang Ma, Fei GuoAAAI 2026
- Difference Vector Equalization for Robust Fine-tuning of Vision-Language ModelsSatoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Taiga Yamane et al.AAAI 2026
- Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action RecognitionHantao Qi, Yan Yan, Junlong Gao, Hanzi WangCVPR 2026
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- Mitigate the Gap: Improving Cross-Modal Alignment in CLIPSedigheh Eslami, Gerard de MeloICLR 2025 · 1 citation
- CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free AttentionZiyu Guo, Renrui Zhang, Longtian Qiu, Xianzheng Ma et al.AAAI 2023 · 182 citations
- I0T: Embedding Standardization Method Towards Zero Modality GapNa Min An, Eunki Kim, James Thorne, Hyunjung ShimACL 2025
- SoftCLIP: Softer Cross-Modal Alignment Makes CLIP StrongerYuting Gao, Jinfeng Liu, Zihan Xu, Tong Wu et al.AAAI 2024 · 80 citations
- Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality InversionMarco Mistretta, Alberto Baldrati, Lorenzo Agnolucci, Marco Bertini et al.ICLR 2025
