Vision Skills Needed to Answer Visual Questions
Xiaoyu Zeng, Yanan Wang, Tai-Yin Chiu, Nilavra Bhattacharya, Danna Gurari
摘要
The task of answering questions about images has garnered attention as a practical service for assisting populations with visual impairments as well as a visual Turing test for the artificial intelligence community. Our first aim is to identify the common vision skills needed for both scenarios. To do so, we analyze the need for four vision skills--object recognition, text recognition, color recognition, and counting--on over 27,000 visual questions from two datasets representing both scenarios. We next quantify the difficulty of these skills for both humans and computers on both datasets. Finally, we propose a novel task of predicting what vision skills are needed to answer a question about an image. Our results reveal (mis)matches between aims of real users of such services and the focus of the AI community. We conclude with a discussion about future directions for addressing the visual question answering task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Grounding Answers for Visual Questions Asked by Visually Impaired PeopleChongyan Chen, Samreen Anjum, Danna GurariCVPR 2022 · 被引用 48 次
- PreSTU: Pre-Training for Scene-Text UnderstandingJihyung Kil, Soravit Changpinyo, Xi Chen, Hexiang Hu 等ICCV 2023 · 被引用 39 次
- Everyday Uncertainty: How Blind People Use GenAI Tools for Information AccessXinru Tang, Ali Abdolrahmani, Darren Gergle, Anne Marie PiperCHI 2025 · 被引用 26 次
- VQA Therapy: Exploring Answer Differences by Visually Grounding AnswersChongyan Chen, Samreen Anjum, Danna GurariICCV 2023 · 被引用 20 次
- "It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language ModelsKapil Garg, Xinru Tang, Jimin Heo, Dwayne R. Morgan 等CHI 2026 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- Right this way: Can VLMs Guide Us to See More to Answer Questions?Li Liu, Diji Yang, Sijia Zhong, Kalyana Suma Sree Tholeti 等NeurIPS 2024 · 被引用 20 次
- Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level VisionHaoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen 等ICLR 2024 · 被引用 258 次
- Exploring Chart Question Answering for Blind and Low Vision UsersJiho Kim, Arjun Srinivasan, Nam Wook Kim, Yea-Seul KimCHI 2023 · 被引用 33 次
- I can't believe there's no images! : Learning Visual Tasks Using Only Language SupervisionSophia Gu, Christopher Clark, Aniruddha KembhaviICCV 2023 · 被引用 41 次
- An Empirical Evaluation of the GPT-4 Multimodal Language Model on Visualization Literacy TasksAlexander Bendeck, John T. StaskoIEEE VIS 2024 · 被引用 40 次
