Grounding Answers for Visual Questions Asked by Visually Impaired People
Chongyan Chen, Samreen Anjum, Danna Gurari
Abstract
Visual question answering is the task of answering questions about images. We introduce the VizWiz-VQA-Grounding dataset, the first dataset that visually grounds answers to visual questions asked by people with visual impairments. We analyze our dataset and compare it with five VQA-Grounding datasets to demonstrate what makes it similar and different. We then evaluate the SOTA VQA and VQA-Grounding models and demonstrate that current SOTA algorithms often fail to identify the correct visual evidence where the answer is located. These models regularly struggle when the visual evidence occupies a small fraction of the image, for images that are higher quality, as well as for visual questions that require skills in text recognition. The dataset, evaluation server, and leaderboard all can be found at the following link: https: //vizwiz.org/tasks-and-datasets/answer- grounding-for-vqa/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5242bb4c-1d4e-4b19-b31c-2f1d17270a2aCited by top-tier papers20
- UNIFIED-IO: A Unified Model for Vision, Language, and Multi-modal TasksJiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi et al.ICLR 2023 · 110 citations
- Can I Trust Your Answer? Visually Grounded Video Question AnsweringJunbin Xiao, Angela Yao, Yicong Li, Tat-Seng ChuaCVPR 2024 · 44 citations
- Everyday Uncertainty: How Blind People Use GenAI Tools for Information AccessXinru Tang, Ali Abdolrahmani, Darren Gergle, Anne Marie PiperCHI 2025 · 26 citations
- Groundhog Grounding Large Language Models to Holistic SegmentationYichi Zhang, Ziqiao Ma, Xiaofeng Gao, Suhaila Shakiah et al.CVPR 2024 · 24 citations
- Right this way: Can VLMs Guide Us to See More to Answer Questions?Li Liu, Diji Yang, Sijia Zhong, Kalyana Suma Sree Tholeti et al.NeurIPS 2024 · 20 citations
Builds on5
- Vision Skills Needed to Answer Visual QuestionsXiaoyu Zeng, Yanan Wang, Tai-Yin Chiu, Nilavra Bhattacharya et al.CSCW 2020 · 17 citations
- Counterfactual VQA: A Cause-Effect Look at Language BiasYulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu et al.CVPR 2021
- Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using CapsulesAisha Urooj Khan, Hilde Kuehne, Kevin Duarte, Chuang Gan et al.CVPR 2021
- Assessing Image Quality Issues for Real-World ProblemsTai-Yin Chiu, Yinan Zhao, Danna GurariCVPR 2020
- "I am uncomfortable sharing what I can't see": Privacy Concerns of the Visually Impaired with Camera Based Assistive ApplicationsTaslima Akter, Bryan Dosono, Tousif Ahmed, Apu Kapadia et al.USENIX Security 2020
Related papers
- VQA Therapy: Exploring Answer Differences by Visually Grounding AnswersChongyan Chen, Samreen Anjum, Danna GurariICCV 2023 · 20 citations
- Acknowledging Focus Ambiguity in Visual QuestionsChongyan Chen, Yu-Yun Tseng, Zhuoheng Li, Anush Venkatesh et al.ICCV 2025 · 1 citation
- Why Does a Visual Question Have Different Answers?Nilavra Bhattacharya, Qing Li, Danna GurariICCV 2019 · 78 citations
- Sentence Attention Blocks for Answer GroundingSeyedalireza Khoshsirat, Chandra KambhamettuICCV 2023 · 8 citations
- CommVQA: Situating Visual Question Answering in Communicative ContextsNandita Naik, Christopher Potts, Elisa KreissEMNLP 2024
