Assessing Image Quality Issues for Real-World Problems
Tai-Yin Chiu, Yinan Zhao, Danna Gurari
Abstract
We introduce a new large-scale dataset that links the assessment of image quality issues to two practical vision tasks: image captioning and visual question answering. First, we identify for 39,181 images taken by people who are blind whether each is sufficient quality to recognize the content as well as what quality flaws are observed from six options. These labels serve as a critical foundation for us to make the following contributions: (1) a new problem and algorithms for deciding whether an image is insufficient quality to recognize the content and so not captionable, (2) a new problem and algorithms for deciding which of six quality flaws an image contains, (3) a new problem and algorithms for deciding whether a visual question is unanswerable due to unrecognizable content versus the content of interest being missing from the field of view, and (4) a novel application of more efficiently creating a large-scale image captioning dataset by automatically deciding whether an image is insufficient quality and so should not be captioned. We publicly-share our datasets and code to facilitate future extensions of this work: https://vizwiz.org .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82aad793-ac96-4d4e-b76d-57f338309952Cited by top-tier papers12
- Grounding Answers for Visual Questions Asked by Visually Impaired PeopleChongyan Chen, Samreen Anjum, Danna GurariCVPR 2022 · 48 citations
- MultiInstruct: Improving Multi-Modal Zero-Shot Learning via Instruction TuningZhiyang Xu, Ying Shen, Lifu HuangACL 2023 · 37 citations
- Disability-First Design and Creation of A Dataset Showing Private Visual Information Collected With People Who Are BlindTanusree Sharma, Abigale Stangl, Lotus Zhang, Yu-Yun Tseng et al.CHI 2023 · 23 citations
- Right this way: Can VLMs Guide Us to See More to Answer Questions?Li Liu, Diji Yang, Sijia Zhong, Kalyana Suma Sree Tholeti et al.NeurIPS 2024 · 20 citations
- Vision Skills Needed to Answer Visual QuestionsXiaoyu Zeng, Yanan Wang, Tai-Yin Chiu, Nilavra Bhattacharya et al.CSCW 2020 · 17 citations
Builds on2
Related papers
- A New Dataset Based on Images Taken by Blind People for Testing the Robustness of Image Classification Models Trained for ImageNet CategoriesReza Akbarian Bafghi, Danna GurariCVPR 2023
- "It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language ModelsKapil Garg, Xinru Tang, Jimin Heo, Dwayne R. Morgan et al.CHI 2026 · 2 citations
- VisAssist: A Visually Impaired-Captured Video Question Answering Benchmark for Assistive SystemsQi Gao, Heng Li, Yixin Zhou, Meixuan Zhou et al.AAAI 2026
- "I Hope This Is Helpful": Understanding Crowdworkers' Challenges and Motivations for an Image Description TaskRachel N. Simons, Danna Gurari, Kenneth R. FleischmannCSCW 2020 · 27 citations
- Accessibility for Color Vision Deficiencies: Challenges and Findings of a Large Scale Study on Paper FiguresKatrin Angerbauer, Nils Rodrigues, René Cutura, Seyda Öney et al.CHI 2022 · 32 citations
