ImageExplorer: Multi-Layered Touch Exploration to Encourage Skepticism Towards Imperfect AI-Generated Image Captions
Jaewook Lee, Jaylin Herskovitz, Yi-Hao Peng, Anhong Guo
Abstract
Blind users rely on alternative text (alt-text) to understand an image; however, alt-text is often missing. AI-generated captions are a more scalable alternative, but they often miss crucial details or are completely incorrect, which users may still falsely trust. In this work, we sought to determine how additional information could help users better judge the correctness of AI-generated captions. We developed ImageExplorer, a touch-based multi-layered image exploration system that allows users to explore the spatial layout and information hierarchies of images, and compared it with popular text-based (Facebook) and touch-based (Seeing AI) image exploration systems in a study with 12 blind participants. We found that exploration was generally successful in encouraging skepticism towards imperfect captions. Moreover, many participants preferred ImageExplorer for its multi-layered and spatial information presentation, and Facebook for its summary and ease of use. Finally, we identify design improvements for effective and explainable image exploration systems for blind users.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07ea7753-b6de-4b81-8481-1902813d4a7eCited by top-tier papers14
- WorldScribe: Towards Context-Aware Live Visual DescriptionsRuei-Che Chang, Yuxuan Liu, Anhong GuoUIST 2024 · 54 citations
- Making Short-Form Videos Accessible with Hierarchical Video SummariesTess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C. Derry et al.CHI 2024 · 37 citations
- SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersZheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen et al.CHI 2024 · 27 citations
- ImageAssist: Tools for Enhancing Touchscreen-Based Image Exploration Systems for Blind and Low Vision UsersVishnu Nair, Hanxiu 'Hazel' Zhu, Brian A. SmithCHI 2023 · 26 citations
- "It's Kind of Context Dependent": Understanding Blind and Low Vision People's Video Accessibility Preferences Across Viewing ScenariosLucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl et al.CHI 2024 · 23 citations
Builds on4
- Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningHarmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana et al.CHI 2020 · 541 citations
- "Person, Shoes, Tree. Is the Person Naked?" What People with Vision Impairments Want in Image DescriptionsAbigale Stangl, Meredith Ringel Morris, Danna GurariCHI 2020 · 136 citations
- Twitter A11y: A Browser Extension to Make Twitter Images AccessibleCole Gleason, Amy Pavel, Emma McCamey, Christina Low et al.CHI 2020 · 123 citations
- Show, Edit and Tell: A Framework for Editing Image CaptionsFawaz Sammani, Luke Melas-KyriaziCVPR 2020
Related papers
- From Provenance to Aberrations: Image Creator and Screen Reader User Perspectives on Alt Text for AI-Generated ImagesMaitraye Das, Alexander J. Fiannaca, Meredith Ringel Morris, Shaun K. Kane et al.CHI 2024 · 32 citations
- AI-Vision: A Three-Layer Accessible Image Exploration System for People with Visual Impairments in ChinaKaixing Zhao, Rui Lai, Bin Guo, Le Liu et al.UbiComp 2024 · 4 citations
- Touch Screen Exploration of Visual Artwork for Blind PeopleDragan Ahmetovic, Nahyun Kwon, Uran Oh, Cristian Bernareggi et al.WWW 2021 · 28 citations
- Accessibility of Profile Pictures: Alt Text and Beyond to Express Identity OnlineMartez E. Mott, John C. Tang, Edward CutrellCHI 2023 · 6 citations
- Communicating Visualizations without Visuals: Investigation of Visualization Alternative Text for People with Visual ImpairmentsCrescentia Jung, Shubham Mehta, Atharva Kulkarni, Yuhang Zhao et al.IEEE VIS 2021 · 2 citations
