How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People
Ricardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y. Lin, Ruiying Hu, Shiri Azenkot
摘要
Multimodal large language models (MLLMs) are changing how Blind and Low Vision (BLV) people access visual information. Unlike traditional visual interpretation tools that only provide descriptions, MLLM-enabled applications offer conversational assistance, where users can ask questions to obtain goal-relevant details. However, evidence about their performance in the real-world and implications for BLV people’s daily lives remains limited. To address this, we conducted a two-week diary study, where we captured 20 BLV participants’ use of an MLLM-enabled visual interpretation application. Although participants rated the visual interpretations of the application as "trustworthy" (mean=3.76 out of 5, max=extremely trustworthy) and "somewhat satisfying" (mean=4.13 out of 5, max=very satisfying), the AI often produced incorrect answers (22.2%) or abstained (10.8%) from responding to users’ requests. Our findings show that while MLLMs can improve visual interpretations’ descriptive accuracy, supporting everyday use also depends on the “visual assistant” skill: behaviors for providing goal-directed, reliable assistance. We conclude by proposing the "visual assistant" skill and guidelines to help MLLM-enabled visual interpretation applications better support BLV people’s access to visual information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable TasksTejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim 等ICLR 2026 · 被引用 154 次
- Improve Vision Language Model Chain-of-thought ReasoningRuohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang 等ACL 2025 · 被引用 135 次
- WorldScribe: Towards Context-Aware Live Visual DescriptionsRuei-Che Chang, Yuxuan Liu, Anhong GuoUIST 2024 · 被引用 54 次
- The Emerging Professional Practice of Remote Sighted Assistance for People with Visual ImpairmentsSooyeon Lee, Madison Reddie, Chun-Hua Tsai, Jordan Beck 等CHI 2020 · 被引用 48 次
相关 Paper
- Beyond Visual Perception: Insights from Smartphone Interaction of Visually Impaired Users with Large Multimodal ModelsJingyi Xie, Rui Yu, He Zhang, Syed Masum Billah 等CHI 2025 · 被引用 40 次
- Investigating Use Cases of AI-Powered Scene Description Applications for Blind and Low Vision PeopleRicardo E. Gonzalez Penuela, Jazmin Collins, Cynthia L. Bennett, Shiri AzenkotCHI 2024 · 被引用 44 次
- Understanding the Use of a Large Language Model-Powered Guide to Make Virtual Reality Accessible for Blind and Low Vision PeopleJazmin Collins, Sharon Y. Lin, Tianqi Liu, Andrea Stevenson Won 等CHI 2026 · 被引用 3 次
- Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired UsersAntonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas 等ACL 2025
- DR-VQA: Decompose-then-Reconstruct for Visual Question Answering in BLV AssistanceBocheng Pan, Hailong Shi, Xingyu GaoACM MM 2025
