ImageExplorer: Multi-Layered Touch Exploration to Encourage Skepticism Towards Imperfect AI-Generated Image Captions
Jaewook Lee, Jaylin Herskovitz, Yi-Hao Peng, Anhong Guo
摘要
Blind users rely on alternative text (alt-text) to understand an image; however, alt-text is often missing. AI-generated captions are a more scalable alternative, but they often miss crucial details or are completely incorrect, which users may still falsely trust. In this work, we sought to determine how additional information could help users better judge the correctness of AI-generated captions. We developed ImageExplorer, a touch-based multi-layered image exploration system that allows users to explore the spatial layout and information hierarchies of images, and compared it with popular text-based (Facebook) and touch-based (Seeing AI) image exploration systems in a study with 12 blind participants. We found that exploration was generally successful in encouraging skepticism towards imperfect captions. Moreover, many participants preferred ImageExplorer for its multi-layered and spatial information presentation, and Facebook for its summary and ease of use. Finally, we identify design improvements for effective and explainable image exploration systems for blind users.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- WorldScribe: Towards Context-Aware Live Visual DescriptionsRuei-Che Chang, Yuxuan Liu, Anhong GuoUIST 2024 · 被引用 54 次
- Making Short-Form Videos Accessible with Hierarchical Video SummariesTess Van Daele, Akhil Iyer, Yuning Zhang, Jalyn C. Derry 等CHI 2024 · 被引用 37 次
- SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersZheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen 等CHI 2024 · 被引用 27 次
- ImageAssist: Tools for Enhancing Touchscreen-Based Image Exploration Systems for Blind and Low Vision UsersVishnu Nair, Hanxiu 'Hazel' Zhu, Brian A. SmithCHI 2023 · 被引用 26 次
- "It's Kind of Context Dependent": Understanding Blind and Low Vision People's Video Accessibility Preferences Across Viewing ScenariosLucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl 等CHI 2024 · 被引用 23 次
它引用的顶会 Paper4
- Interpreting Interpretability: Understanding Data Scientists' Use of Interpretability Tools for Machine LearningHarmanpreet Kaur, Harsha Nori, Samuel Jenkins, Rich Caruana 等CHI 2020 · 被引用 541 次
- "Person, Shoes, Tree. Is the Person Naked?" What People with Vision Impairments Want in Image DescriptionsAbigale Stangl, Meredith Ringel Morris, Danna GurariCHI 2020 · 被引用 136 次
- Twitter A11y: A Browser Extension to Make Twitter Images AccessibleCole Gleason, Amy Pavel, Emma McCamey, Christina Low 等CHI 2020 · 被引用 123 次
- Show, Edit and Tell: A Framework for Editing Image CaptionsFawaz Sammani, Luke Melas-KyriaziCVPR 2020
相关 Paper
- From Provenance to Aberrations: Image Creator and Screen Reader User Perspectives on Alt Text for AI-Generated ImagesMaitraye Das, Alexander J. Fiannaca, Meredith Ringel Morris, Shaun K. Kane 等CHI 2024 · 被引用 32 次
- AI-Vision: A Three-Layer Accessible Image Exploration System for People with Visual Impairments in ChinaKaixing Zhao, Rui Lai, Bin Guo, Le Liu 等UbiComp 2024 · 被引用 4 次
- Touch Screen Exploration of Visual Artwork for Blind PeopleDragan Ahmetovic, Nahyun Kwon, Uran Oh, Cristian Bernareggi 等WWW 2021 · 被引用 28 次
- Accessibility of Profile Pictures: Alt Text and Beyond to Express Identity OnlineMartez E. Mott, John C. Tang, Edward CutrellCHI 2023 · 被引用 6 次
- Communicating Visualizations without Visuals: Investigation of Visualization Alternative Text for People with Visual ImpairmentsCrescentia Jung, Shubham Mehta, Atharva Kulkarni, Yuhang Zhao 等IEEE VIS 2021 · 被引用 2 次
