Grounding Answers for Visual Questions Asked by Visually Impaired People
Chongyan Chen, Samreen Anjum, Danna Gurari
摘要
Visual question answering is the task of answering questions about images. We introduce the VizWiz-VQA-Grounding dataset, the first dataset that visually grounds answers to visual questions asked by people with visual impairments. We analyze our dataset and compare it with five VQA-Grounding datasets to demonstrate what makes it similar and different. We then evaluate the SOTA VQA and VQA-Grounding models and demonstrate that current SOTA algorithms often fail to identify the correct visual evidence where the answer is located. These models regularly struggle when the visual evidence occupies a small fraction of the image, for images that are higher quality, as well as for visual questions that require skills in text recognition. The dataset, evaluation server, and leaderboard all can be found at the following link: https: //vizwiz.org/tasks-and-datasets/answer- grounding-for-vqa/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- UNIFIED-IO: A Unified Model for Vision, Language, and Multi-modal TasksJiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mottaghi 等ICLR 2023 · 被引用 110 次
- Can I Trust Your Answer? Visually Grounded Video Question AnsweringJunbin Xiao, Angela Yao, Yicong Li, Tat-Seng ChuaCVPR 2024 · 被引用 44 次
- Everyday Uncertainty: How Blind People Use GenAI Tools for Information AccessXinru Tang, Ali Abdolrahmani, Darren Gergle, Anne Marie PiperCHI 2025 · 被引用 26 次
- Groundhog Grounding Large Language Models to Holistic SegmentationYichi Zhang, Ziqiao Ma, Xiaofeng Gao, Suhaila Shakiah 等CVPR 2024 · 被引用 24 次
- Right this way: Can VLMs Guide Us to See More to Answer Questions?Li Liu, Diji Yang, Sijia Zhong, Kalyana Suma Sree Tholeti 等NeurIPS 2024 · 被引用 20 次
它引用的顶会 Paper5
- Vision Skills Needed to Answer Visual QuestionsXiaoyu Zeng, Yanan Wang, Tai-Yin Chiu, Nilavra Bhattacharya 等CSCW 2020 · 被引用 17 次
- Counterfactual VQA: A Cause-Effect Look at Language BiasYulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu 等CVPR 2021
- Found a Reason for me? Weakly-supervised Grounded Visual Question Answering using CapsulesAisha Urooj Khan, Hilde Kuehne, Kevin Duarte, Chuang Gan 等CVPR 2021
- Assessing Image Quality Issues for Real-World ProblemsTai-Yin Chiu, Yinan Zhao, Danna GurariCVPR 2020
- "I am uncomfortable sharing what I can't see": Privacy Concerns of the Visually Impaired with Camera Based Assistive ApplicationsTaslima Akter, Bryan Dosono, Tousif Ahmed, Apu Kapadia 等USENIX Security 2020
相关 Paper
- VQA Therapy: Exploring Answer Differences by Visually Grounding AnswersChongyan Chen, Samreen Anjum, Danna GurariICCV 2023 · 被引用 20 次
- Acknowledging Focus Ambiguity in Visual QuestionsChongyan Chen, Yu-Yun Tseng, Zhuoheng Li, Anush Venkatesh 等ICCV 2025 · 被引用 1 次
- Why Does a Visual Question Have Different Answers?Nilavra Bhattacharya, Qing Li, Danna GurariICCV 2019 · 被引用 78 次
- Sentence Attention Blocks for Answer GroundingSeyedalireza Khoshsirat, Chandra KambhamettuICCV 2023 · 被引用 8 次
- CommVQA: Situating Visual Question Answering in Communicative ContextsNandita Naik, Christopher Potts, Elisa KreissEMNLP 2024
