Post-hoc Probabilistic Vision-Language Models
Anton Baumann, Rui Li, Marcus Klasson, Santeri Mentu, Shyamgopal Karthik, Zeynep Akata, Arno Solin, Martin Trapp
Abstract
Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs deterministically map images and text descriptions to a joint latent space in which their similarity is assessed using the cosine similarity. However, a deterministic mapping of inputs fails to capture uncertainties over concepts arising from domain shifts when used in downstream tasks. In this work, we propose post-hoc uncertainty estimation in VLMs that does not require additional training. Our method leverages a Bayesian posterior approximation over the last layers in VLMs and analytically quantifies uncertainties over cosine similarities. We demonstrate its effectiveness for uncertainty quantification and support set selection in active learning. Compared to baselines, we obtain improved and well-calibrated predictive uncertainties, interpretable uncertainty estimates, and sample-efficient active learning. Our results show promise for safety-critical applications of large-scale models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- InfoNCE Induces Gaussian DistributionRoy Betser, Eyal Gofer, Meir Yossef Levi, Guy GilboaICLR 2026 · 17 citations
- Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit HypersphereLi Ju, Max Andersson, Stina Fredriksson, Edward Glöckner et al.NeurIPS 2025 · 5 citations
- Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient TransformersFiras Gabetni, Giuseppe Curci, Andrea Pilzer, Subhankar Roy et al.ICLR 2026 · 5 citations
- Epistemic Uncertainty Quantification for Pre-trained VLMs via Riemannian Flow MatchingLi Ju, Mayank Nautiyal, Andreas Hellander, Ekta Vats et al.ICML 2026 · 2 citations
- Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active LearningZhuofan Xie, Zishan Lin, Jinliang Lin, Jie Qi et al.CVPR 2026 · 1 citation
Builds on38
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
Related papers
- ProbVLM: Probabilistic Adapter for Frozen Vison-Language ModelsUddeshya Upadhyay, Shyamgopal Karthik, Massimiliano Mancini, Zeynep AkataICCV 2023 · 41 citations
- ViLU: Learning Vision-Language Uncertainties for Failure PredictionMarc Lafon, Yannis Karmim, Julio Silva-Rodríguez, Paul Couairon et al.ICCV 2025 · 1 citation
- Probabilistic Language-Image Pre-TrainingSanghyuk Chun, Wonjae Kim, Song Park, Sangdoo YunICLR 2025
- MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image SegmentationTaha Koleilat, Hojat Asgariandehkordi, Omid Nejatimanzari, Berardino Barile et al.CVPR 2026 · 4 citations
- BLIPs: Bayesian Learned Interatomic PotentialsDario Coscia, Pim de Haan, Max WellingICML 2026 · 6 citations
