Whitened CLIP as a Likelihood Surrogate of Images and Captions
Roy Betser, Meir Yossef Levi, Guy Gilboa
Abstract
Likelihood approximations for images are not trivial to compute and can be useful in many applications. We examine the use of Contrastive Language-Image Pre-training (CLIP) to assess the likelihood of images and captions. We introduce Whitened CLIP, a novel transformation of the CLIP latent space via an invertible linear operation. This transformation ensures that each feature in the embedding space has zero mean, unit standard deviation, and no correlation with all other features, resulting in an identity covariance matrix. We show that the whitened embedding statistics can be well approximated by a standard normal distribution, allowing log-likelihood to be estimated using the squared Euclidean norm in the whitened space. The whitening procedure is completely training-free and uses a precomputed whitening matrix, making it extremely fast. We present several preliminary experiments demonstrating the properties and applicability of these likelihood scores to images and captions. Our code is available at github.com/rbetser/W_CLIP/tree/main.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7bfa0920-bdb3-4680-9a72-e7b758c1a061Cited by top-tier papers7
- InfoNCE Induces Gaussian DistributionRoy Betser, Eyal Gofer, Meir Yossef Levi, Guy GilboaICLR 2026 · 17 citations
- Training-free Detection of Generated Videos via Spatial-Temporal LikelihoodsOmer Ben Hayun, Roy Betser, Meir Yossef Levi, Levi Kassel et al.CVPR 2026 · 7 citations
- The Universal Normal EmbeddingChen Tasker, Roy Betser, Eyal Gofer, Meir Yossef Levi et al.CVPR 2026 · 4 citations
- The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-Modal DivergenceYichao Cai, Zhen Zhang, Yuhang Liu, Javen Qinfeng ShiICML 2026 · 2 citations
- Epistemic Uncertainty Quantification for Pre-trained VLMs via Riemannian Flow MatchingLi Ju, Mayank Nautiyal, Andreas Hellander, Ekta Vats et al.ICML 2026 · 2 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
Related papers
- Learning Invariant Causal Mechanism from Vision-Language ModelsZeen Song, Siyu Zhao, Xingyu Zhang, Jiangmeng Li et al.ICML 2025
- The Double-Ellipsoid Geometry of CLIPMeir Yossef Levi, Guy GilboaICML 2025
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 1,208 citations
- Modeling Caption Diversity in Contrastive Vision-Language PretrainingSamuel Lavoie, Polina Kirichenko, Mark Ibrahim, Mido Assran et al.ICML 2024 · 44 citations
- Iterative Prompt Learning for Unsupervised Backlit Image EnhancementZhexin Liang, Chongyi Li, Shangchen Zhou, Ruicheng Feng et al.ICCV 2023 · 196 citations
