Acknowledging Focus Ambiguity in Visual Questions
Chongyan Chen, Yu-Yun Tseng, Zhuoheng Li, Anush Venkatesh, Danna Gurari
Abstract
No published work on visual question answering (VQA) accounts for ambiguity regarding where the content described in the question is located in the image. To fill this gap, we introduce VQ-FocusAmbiguity, the first VQA dataset that visually grounds each plausible image region a question could refer to when arriving at valid answers. We next analyze and compare our dataset to existing datasets to reveal its unique properties. Finally, we benchmark modern models for two novel tasks related to acknowledging focus ambiguity: recognizing whether a visual question has focus ambiguity and locating all plausible focus regions within the image. Results show that the dataset is challenging for modern models. To facilitate future progress on these tasks, we publicly share the dataset with an evaluation server at https://vizwiz.org/tasks-and-datasets/ focus-ambiguity-in-visual-questions/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08b6d296-d3d6-487b-95f9-4c0ea1d3f9caCited by top-tier papers2
- AQuA: Toward Strategic Response Generation for Ambiguous Visual QuestionsJihyoung Jang, Hyounghun KimICLR 2026 · 1 citation
- MMBench-Live: A Continuously Evolving Benchmark for Multimodal ModelsYuanzhi Liu, Shousheng Zhao, Bo Zhou, Kongming Liang et al.ICML 2026
Builds on20
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Multimodal Few-Shot Learning with Frozen Language ModelsMaria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami et al.NeurIPS 2021 · 1,020 citations
Related papers
- Grounding Answers for Visual Questions Asked by Visually Impaired PeopleChongyan Chen, Samreen Anjum, Danna GurariCVPR 2022 · 48 citations
- VQA Therapy: Exploring Answer Differences by Visually Grounding AnswersChongyan Chen, Samreen Anjum, Danna GurariICCV 2023 · 20 citations
- Looking Beyond the One: Operationalizing and Eliciting Visual Ambiguity in VLLMsYuchong Chen, Bowei Zou, Yuhan Chen, Yifan Fan et al.ACL 2026
- On the General Value of Evidence, and Bilingual Scene-Text Visual Question AnsweringXinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng et al.CVPR 2020
- Why Does a Visual Question Have Different Answers?Nilavra Bhattacharya, Qing Li, Danna GurariICCV 2019 · 78 citations
