RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs
Alberto Testoni, Barbara Plank, Raquel Fernández
摘要
Ambiguity resolution is key to effective communication. While humans effortlessly address ambiguity through conversational grounding strategies, the extent to which current language models can emulate these strategies remains unclear. In this work, we examine referential ambiguity in image-based question answering by introducing RACQUET, a carefully curated dataset targeting distinct aspects of ambiguity. Through a series of evaluations, we reveal significant limitations and problems of overconfidence of state-of-the-art large multimodal language models in addressing ambiguity in their responses. The overconfidence issue becomes particularly relevant for RACQUET-BIAS, a subset designed to analyze a critical yet underexplored problem: failing to address ambiguity leads to stereotypical, socially biased responses. Our results underscore the urgency of equipping models with robust strategies to deal with uncertainty without resorting to undesirable stereotypes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Looking Beyond the One: Operationalizing and Eliciting Visual Ambiguity in VLLMsYuchong Chen, Bowei Zou, Yuhan Chen, Yifan Fan 等ACL 2026
- Hidden in Plain Sight: Reasoning in Underspecified and Misspecified Scenarios for Multimodal LLMsQianqi Yan, Hongquan Li, Shan Jiang, Yang Zhao 等EMNLP 2025
- Using Perspectival Words Is Harder Than Vocabulary Words for Humans - and Even More So for Multimodal Language ModelsDota Tianai Dong, Yifan Luo, Po-Ya Angela Wang, Asli Özyürek 等ACL 2026
它引用的顶会 Paper7
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 被引用 162 次
- Why Does a Visual Question Have Different Answers?Nilavra Bhattacharya, Qing Li, Danna GurariICCV 2019 · 被引用 78 次
- We're Afraid Language Models Aren't Modeling AmbiguityAlisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr 等EMNLP 2023 · 被引用 35 次
- Dealing with Semantic Underspecification in Multimodal NLPSandro PezzelleACL 2023 · 被引用 5 次
- Why Did the Chicken Cross the Road? Rephrasing and Analyzing Ambiguous Questions in VQAElias Stengel-Eskin, Jimena Guallar-Blasco, Yi Zhou, Benjamin Van DurmeACL 2023 · 被引用 4 次
相关 Paper
- AmbiRefer3D: 3D Visual Grounding with Referential AmbiguityRongjiang Zhu, Wei Kang, Zeqi Liu, Chen junyu 等ICML 2026
- VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language ModelsChahat Raj, Bowen Wei, Aylin Caliskan, Antonios Anastasopoulos 等ACL 2026 · 被引用 3 次
- VAGUE: Visual Contexts Clarify Ambiguous ExpressionsHeejeong Nam, Jinwoo Ahn, Keummin Ka, Jiwan Chung 等ICCV 2025 · 被引用 1 次
- Acknowledging Focus Ambiguity in Visual QuestionsChongyan Chen, Yu-Yun Tseng, Zhuoheng Li, Anush Venkatesh 等ICCV 2025 · 被引用 1 次
- Resolving Ambiguities in Text-to-Image Generative ModelsNinareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala 等ACL 2023 · 被引用 8 次
