Test-Time Distribution Normalization for Contrastively Learned Visual-language Models
Yifei Zhou, Juntao Ren, Fengyu Li, Ramin Zabih, Ser Nam Lim
Abstract
Advances in the field of vision-language contrastive learning have made it possible for many downstream applications to be carried out efficiently and accurately by simply taking the dot product between image and text representations. One of the most representative approaches proposed recently known as CLIP [50] has garnered widespread adoption due to its effectiveness. CLIP is trained with an InfoNCE loss that takes into account both positive and negative samples to help learn a much more robust representation space. This paper reveals that the common downstream practice of taking a dot product is only a zeroth-order approximation of the optimization goal, resulting in a loss of information during test-time. Intuitively, since the model has been optimized based on the InfoNCE loss, test-time procedures should also be in alignment. The question lies in how one can retrieve any semblance of negative samples information during inference in a computationally efficient way. To this end, we propose Distribution Normalization (DN), where we approximate the mean representation of a batch of test samples and use such a mean to represent what would be analogous to negative samples in the InfoNCE loss. DN requires no retraining or fine-tuning and can be effortlessly applied during inference. Extensive experiments on a wide variety of downstream tasks exhibit a clear advantage of DN over the dot product on top of other existing test-time augmentation methods. Our code is available at https://github.com/fengyuli2002/distribution-normalization .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec9fb77e-17c9-4f06-a880-43ffebdd10a0Cited by top-tier papers6
- Test-Time Retrieval-Augmented Adaptation for Vision-Language ModelsXinqi Fan, Xueli Chen, Luoxiao Yang, Chuin Hong Yap et al.ICCV 2025 · 4 citations
- Inverse Optimal Transport for Efficient Adaptation of Vision-Language ModelsShupeng Qiu, Chuan-Xian RenAAAI 2026 · 1 citation
- Efficient and Context-Aware Label Propagation for Zero-/Few-Shot Training-Free Adaptation of Vision-Language ModelYushu Li, Yongyi Su, Adam Goodge, Kui Jia et al.ICLR 2025
- Compositional Caching for Training-free Open-vocabulary Attribute DetectionMarco Garosi, Alessandro Conti, Gaowen Liu, Elisa Ricci et al.CVPR 2025
- Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language ModelsYabin Zhang, Wenjie Zhu, Hui Tang, Zhiyuan Ma et al.CVPR 2024
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- WaveDN: A Wavelet-based Training-free Zero-shot Enhancement for Vision-Language ModelsJiulin Li, Mengyu Yang, Ye Tian, Lanshan Zhang et al.ACM MM 2024 · 3 citations
- Towards Robustness Prompt Tuning with Fully Test-Time Adaptation for CLIP's Zero-Shot GeneralizationRan Wang, Hua Zuo, Zhen Fang, Jie LuACM MM 2024 · 7 citations
- Data Efficient Language-Supervised Zero-Shot Recognition with Optimal Transport DistillationBichen Wu, Ruizhe Cheng, Peizhao Zhang, Tianren Gao et al.ICLR 2022 · 57 citations
- DOTA: Distributional Test-time Adaptation of Vision-Language ModelsZongbo Han, Jialong Yang, Guangyu Wang, Junfan Li et al.NeurIPS 2025 · 25 citations
- Open-Vocabulary Customization from CLIP via Data-Free Knowledge DistillationYongxian Wei, Zixuan Hu, Li Shen, Zhenyi Wang et al.ICLR 2025
