Whitened CLIP as a Likelihood Surrogate of Images and Captions
Roy Betser, Meir Yossef Levi, Guy Gilboa
摘要
Likelihood approximations for images are not trivial to compute and can be useful in many applications. We examine the use of Contrastive Language-Image Pre-training (CLIP) to assess the likelihood of images and captions. We introduce Whitened CLIP, a novel transformation of the CLIP latent space via an invertible linear operation. This transformation ensures that each feature in the embedding space has zero mean, unit standard deviation, and no correlation with all other features, resulting in an identity covariance matrix. We show that the whitened embedding statistics can be well approximated by a standard normal distribution, allowing log-likelihood to be estimated using the squared Euclidean norm in the whitened space. The whitening procedure is completely training-free and uses a precomputed whitening matrix, making it extremely fast. We present several preliminary experiments demonstrating the properties and applicability of these likelihood scores to images and captions. Our code is available at github.com/rbetser/W_CLIP/tree/main.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- InfoNCE Induces Gaussian DistributionRoy Betser, Eyal Gofer, Meir Yossef Levi, Guy GilboaICLR 2026 · 被引用 17 次
- Training-free Detection of Generated Videos via Spatial-Temporal LikelihoodsOmer Ben Hayun, Roy Betser, Meir Yossef Levi, Levi Kassel 等CVPR 2026 · 被引用 7 次
- The Universal Normal EmbeddingChen Tasker, Roy Betser, Eyal Gofer, Meir Yossef Levi 等CVPR 2026 · 被引用 4 次
- The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-Modal DivergenceYichao Cai, Zhen Zhang, Yuhang Liu, Javen Qinfeng ShiICML 2026 · 被引用 2 次
- Epistemic Uncertainty Quantification for Pre-trained VLMs via Riemannian Flow MatchingLi Ju, Mayank Nautiyal, Andreas Hellander, Ekta Vats 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
相关 Paper
- Learning Invariant Causal Mechanism from Vision-Language ModelsZeen Song, Siyu Zhao, Xingyu Zhang, Jiangmeng Li 等ICML 2025
- The Double-Ellipsoid Geometry of CLIPMeir Yossef Levi, Guy GilboaICML 2025
- Exploring CLIP for Assessing the Look and Feel of ImagesJianyi Wang, Kelvin C. K. Chan, Chen Change LoyAAAI 2023 · 被引用 1,208 次
- Modeling Caption Diversity in Contrastive Vision-Language PretrainingSamuel Lavoie, Polina Kirichenko, Mark Ibrahim, Mido Assran 等ICML 2024 · 被引用 44 次
- Iterative Prompt Learning for Unsupervised Backlit Image EnhancementZhexin Liang, Chongyi Li, Shangchen Zhou, Ruicheng Feng 等ICCV 2023 · 被引用 196 次
