Looking Beyond the One: Operationalizing and Eliciting Visual Ambiguity in VLLMs
Yuchong Chen, Bowei Zou, Yuhan Chen, Yifan Fan, Xinyu Li, Shujun Cao, Yu Hong
摘要
Visual questions are often ambiguous: the same image-question pair may admit multiple valid answers depending on which region is referenced. However, current Visual Question Answering (VQA) systems typically collapse this ambiguity, committing to a single interpretation during decoding and evaluation. In this work, we study visual question ambiguity from a grounded, region-centric perspective. We operationalize ambiguity as the existence of multiple distinct answer-supporting regions in an image, each independently yielding a valid answer. This formulation makes ambiguity observable without requiring exhaustive multi-answer annotations. Based on this definition, we conduct a systematic empirical study of state-of-the-art Visual Large Language Models (VLLMs). We find that, under default decoding, VLLMs consistently under-report ambiguity-even when multiple valid visual groundings are present. Importantly, probing model hidden states reveals that ambiguity-related signals are already encoded in their internal representations, despite not being reliably expressed in outputs. Finally, we show that selectively activating multi-focus answering based on these signals can recover additional valid answers while avoiding excessive hallucination. Together, our results suggest that ambiguity in VQA is not merely an annotation artifact or capability limitation, but a property that VLLMs internally recognize yet often fail to surface under standard decoding assumptions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 被引用 162 次
- Enabling Large Language Models to Generate Text with CitationsTianyu Gao, Howard Yen, Jiatong Yu, Danqi ChenEMNLP 2023 · 被引用 152 次
- ASQA: Factoid Questions Meet Long-Form AnswersIvan Stelmakh, Yi Luan, Bhuwan Dhingra, Ming-Wei ChangEMNLP 2022 · 被引用 51 次
- Getting a CLUE: A Method for Explaining Uncertainty EstimatesJavier Antorán, Umang Bhatt, Tameem Adel, Adrian Weller 等ICLR 2021 · 被引用 41 次
- VQA Therapy: Exploring Answer Differences by Visually Grounding AnswersChongyan Chen, Samreen Anjum, Danna GurariICCV 2023 · 被引用 20 次
相关 Paper
- AVAM: A Universal Training-Free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-Image Question AnsweringKang Zeng, Guojin Zhong, Jintao Cheng, Jin Yuan 等ICCV 2025 · 被引用 1 次
- Acknowledging Focus Ambiguity in Visual QuestionsChongyan Chen, Yu-Yun Tseng, Zhuoheng Li, Anush Venkatesh 等ICCV 2025 · 被引用 1 次
- Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language ModelsRuiying Peng, Xueyu Wu, Jing Lei, Lu Hou 等CVPR 2026 · 被引用 4 次
- Unveiling the Response of Large Vision-Language Models to Visually Absent TokensSohee Kim, Soohyun Ryu, Joonhyung Park, Eunho YangEMNLP 2025
- Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?Gregor Geigle, Radu Timofte, Goran GlavasEMNLP 2024 · 被引用 1 次
