ViLU: Learning Vision-Language Uncertainties for Failure Prediction
Marc Lafon, Yannis Karmim, Julio Silva-Rodríguez, Paul Couairon, Clément Rambour, Raphaël Fournier-S'niehotta, Ismail Ben Ayed, Jose Dolz, Nicolas Thome
摘要
Reliable Uncertainty Quantification (UQ) and failure prediction remain open challenges for Vision-Language Models (VLMs). We introduce ViLU, a new Vision-Language Uncertainty quantification framework that contextualizes uncertainty estimates by leveraging all task-relevant textual representations. ViLU constructs an uncertainty-aware multi-modal representation by integrating the visual embedding, the predicted textual embedding, and an image-conditioned textual representation via cross-attention. Unlike traditional UQ methods based on loss prediction, ViLU trains an uncertainty predictor as a binary classifier to distinguish correct from incorrect predictions using a weighted binary cross-entropy loss, making it loss-agnostic. In particular, our proposed approach is well-suited for post-hoc settings, where only vision and text embeddings are available without direct access to the model itself. Extensive experiments on diverse datasets show the significant gains of our method compared to state-of-the-art failure prediction methods. We apply our method to standard classification datasets, such as ImageNet-1k, as well as large-scale image-caption datasets like CC12M and LAION-400M. Ablation studies highlight the critical role of our architecture and training in achieving effective uncertainty quantification. Our code is publicly available and can be found here: https://github.com/ykrmm/ViLU.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Post-hoc Probabilistic Vision-Language ModelsAnton Baumann, Rui Li, Marcus Klasson, Santeri Mentu 等ICLR 2026 · 被引用 14 次
- Ensembling Pruned Attention Heads For Uncertainty-Aware Efficient TransformersFiras Gabetni, Giuseppe Curci, Andrea Pilzer, Subhankar Roy 等ICLR 2026 · 被引用 5 次
- Revisiting Confidence Calibration for Misclassification Detection in VLMsJincheng Huang, Jie Xu, Xiaoshuang Shi, Ping Hu 等ICLR 2026
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- Revisiting the Calibration of Modern Neural NetworksMatthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis 等NeurIPS 2021 · 被引用 633 次
相关 Paper
- Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit HypersphereLi Ju, Max Andersson, Stina Fredriksson, Edward Glöckner 等NeurIPS 2025 · 被引用 5 次
- Probabilistic Language-Image Pre-TrainingSanghyuk Chun, Wonjae Kim, Song Park, Sangdoo YunICLR 2025
- Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty QuantificationTao Huang, Rui Wang, Xiaofei Liu, Yi Qin 等ICLR 2026 · 被引用 4 次
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang 等NeurIPS 2025 · 被引用 7 次
- MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image SegmentationTaha Koleilat, Hojat Asgariandehkordi, Omid Nejatimanzari, Berardino Barile 等CVPR 2026 · 被引用 4 次
