ViLU: Learning Vision-Language Uncertainties for Failure Prediction
Marc Lafon, Yannis Karmim, Julio Silva-Rodríguez, Paul Couairon, Clément Rambour, Raphaël Fournier-S'niehotta, Ismail Ben Ayed, Jose Dolz, Nicolas Thome
Abstract
Reliable Uncertainty Quantification (UQ) and failure prediction remain open challenges for Vision-Language Models (VLMs). We introduce ViLU, a new Vision-Language Uncertainty quantification framework that contextualizes uncertainty estimates by leveraging all task-relevant textual representations. ViLU constructs an uncertainty-aware multi-modal representation by integrating the visual embedding, the predicted textual embedding, and an image-conditioned textual representation via cross-attention. Unlike traditional UQ methods based on loss prediction, ViLU trains an uncertainty predictor as a binary classifier to distinguish correct from incorrect predictions using a weighted binary cross-entropy loss, making it loss-agnostic. In particular, our proposed approach is well-suited for post-hoc settings, where only vision and text embeddings are available without direct access to the model itself. Extensive experiments on diverse datasets show the significant gains of our method compared to state-of-the-art failure prediction methods. We apply our method to standard classification datasets, such as ImageNet-1k, as well as large-scale image-caption datasets like CC12M and LAION-400M. Ablation studies highlight the critical role of our architecture and training in achieving effective uncertainty quantification. Our code is publicly available and can be found here: https://github.com/ykrmm/ViLU.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b2609bf-69c4-4eaf-bd8c-4063a26cac25Cited by top-tier papers3
- Post-hoc Probabilistic Vision-Language ModelsAnton Baumann, Rui Li, Marcus Klasson, Santeri Mentu et al.ICLR 2026 · 14 citations
- Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient TransformersFiras Gabetni, Giuseppe Curci, Andrea Pilzer, Subhankar Roy et al.ICLR 2026 · 5 citations
- Revisiting Confidence Calibration for Misclassification Detection in VLMsJincheng Huang, Jie Xu, Xiaoshuang Shi, Ping Hu et al.ICLR 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis et al.NeurIPS 2021 · 633 citations
Related papers
- Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit HypersphereLi Ju, Max Andersson, Stina Fredriksson, Edward Glöckner et al.NeurIPS 2025 · 5 citations
- Probabilistic Language-Image Pre-TrainingSanghyuk Chun, Wonjae Kim, Song Park, Sangdoo YunICLR 2025
- Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty QuantificationTao Huang, Rui Wang, Xiaofei Liu, Yi Qin et al.ICLR 2026 · 4 citations
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang et al.NeurIPS 2025 · 7 citations
- MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image SegmentationTaha Koleilat, Hojat Asgariandehkordi, Omid Nejatimanzari, Berardino Barile et al.CVPR 2026 · 4 citations
