Seeing is Believing, but How Much? A Comprehensive Analysis of Verbalized Calibration in Vision-Language Models
Weihao Xuan, Qingcheng Zeng, Heli Qi, Junjue Wang, Naoto Yokoya
Abstract
Uncertainty quantification is essential for assessing the reliability and trustworthiness of modern AI systems. Among existing approaches, verbalized uncertainty-where models express their confidence through natural language-has emerged as a lightweight and interpretable solution in large language models (LLMs). However, its effectiveness in vision-language models (VLMs) remains insufficiently studied. In this work, we conduct a comprehensive evaluation of verbalized confidence in VLMs, spanning three model categories, four task domains, and three evaluation scenarios. Our results show that current VLMs often display notable miscalibration across diverse tasks and settings. Notably, visual reasoning models (i.e., thinking with images) consistently exhibit better calibration, suggesting that modality-specific reasoning is critical for reliable uncertainty estimation. To further address calibration challenges, we introduce VI-SUAL CONFIDENCE-AWARE PROMPTING, a two-stage prompting strategy that improves confidence alignment in multimodal settings. Overall, our study highlights the inherent miscalibration in VLMs across modalities. More broadly, our findings underscore the fundamental importance of modality alignment and model faithfulness in advancing reliable multimodal systems. General Setting (Image + Text Instruction) Q: Given the image, choose an integral expression that can be used to find the area of R.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d6f88bf-ac65-4801-90e9-0eeac21e4275Cited by top-tier papers4
- The Confidence Dichotomy: Analyzing and Mitigating Miscalibration in Tool-Use AgentsWeihao Xuan, Qingcheng Zeng, Heli Qi, Yunze Xiao et al.ACL 2026 · 4 citations
- Thinking Out Loud: Do Reasoning Models Know When They're Right?Qingcheng Zeng, Weihao Xuan, Leyang Cui, Rob VoigtEMNLP 2025 · 1 citation
- Dual-Level Confidence based Implicit Self-Refinement for Medical Visual Question AnsweringMeihong Pan, Yefeng ZhengCVPR 2026
- Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model ExplanationsKeyu He, Tejas Srinivasan, Brihi Joshi, Xiang Ren et al.ACL 2026
Builds on8
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li et al.ICLR 2024 · 867 citations
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang et al.ACL 2025 · 377 citations
- Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic SimilaritiesAlexander Nikitin, Jannik Kossen, Yarin Gal, Pekka MarttinenNeurIPS 2024 · 197 citations
Related papers
- VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models ReasoningWenyi Xiao, Xinchi Xu, Leilei GanACL 2026 · 2 citations
- Confidence is Not Universal: Task-Dependent Calibration and Emergent Behavior in LLMsChaeyun Jang, Moonseok Choi, Yegon Kim, Seungyoo Lee et al.ICML 2026
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMsGabrielle Kaili-May Liu, Gal Yona, Avi Caciularu, Idan Szpektor et al.EMNLP 2025
- An Empirical Study Into What Matters for Calibrating Vision-Language ModelsWeijie Tu, Weijian Deng, Dylan Campbell, Stephen Gould et al.ICML 2024 · 18 citations
- Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question AnsweringMahsa Mozaffari, Hitesh Sapkota, Xumin Liu, Qi YuICLR 2026
