A New Dataset Based on Images Taken by Blind People for Testing the Robustness of Image Classification Models Trained for ImageNet Categories
Reza Akbarian Bafghi, Danna Gurari
Abstract
Our goal is to improve upon the status quo for designing image classification models trained in one domain that perform well on images from another domain. Complementing existing work in robustness testing, we introduce the first dataset for this purpose which comes from an authentic use case where photographers wanted to learn about the content in their images. We built a new test set using 8,900 images taken by people who are blind for which we collected metadata to indicate the presence versus absence of 200 ImageNet object categories. We call this dataset VizWiz-Classification. We characterize this dataset and how it compares to the mainstream datasets for evaluating how well ImageNet-trained classification models generalize. Finally, we analyze the performance of 100 Ima-geNet classification models on our new test dataset. Our fine-grained analysis demonstrates that these models struggle on images with quality issues. To enable future extensions to this work, we share our new dataset with evaluation server at: https://vizwiz.org/tasks-and- datasets/image-classification.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b39c4850-7879-4ddc-89dc-5a883aff86d3Cited by top-tier papers2
- LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic ImagesJing Zhang, Irving Fang, Hao Wu, Akshat Kaushik et al.CVPR 2024 · 5 citations
- Explaining CLIP's Performance Disparities on Data from Blind/Low Vision UsersDaniela Massiceti, Camilla Longden, Agnieszka Slowik, Samuel Wills et al.CVPR 2024 · 3 citations
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
Related papers
- Assessing Image Quality Issues for Real-World ProblemsTai-Yin Chiu, Yinan Zhao, Danna GurariCVPR 2020
- Grounding Answers for Visual Questions Asked by Visually Impaired PeopleChongyan Chen, Samreen Anjum, Danna GurariCVPR 2022 · 48 citations
- ORBIT: A Real-World Few-Shot Dataset for Teachable Object RecognitionDaniela Massiceti, Luisa M. Zintgraf, John Bronskill, Lida Theodorou et al.ICCV 2021 · 55 citations
- VisAssist: A Visually Impaired-Captured Video Question Answering Benchmark for Assistive SystemsQi Gao, Heng Li, Yixin Zhou, Meixuan Zhou et al.AAAI 2026
- Why Does a Visual Question Have Different Answers?Nilavra Bhattacharya, Qing Li, Danna GurariICCV 2019 · 78 citations
