How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision People
Ricardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y. Lin, Ruiying Hu, Shiri Azenkot
Abstract
Multimodal large language models (MLLMs) are changing how Blind and Low Vision (BLV) people access visual information. Unlike traditional visual interpretation tools that only provide descriptions, MLLM-enabled applications offer conversational assistance, where users can ask questions to obtain goal-relevant details. However, evidence about their performance in the real-world and implications for BLV people’s daily lives remains limited. To address this, we conducted a two-week diary study, where we captured 20 BLV participants’ use of an MLLM-enabled visual interpretation application. Although participants rated the visual interpretations of the application as "trustworthy" (mean=3.76 out of 5, max=extremely trustworthy) and "somewhat satisfying" (mean=4.13 out of 5, max=very satisfying), the AI often produced incorrect answers (22.2%) or abstained (10.8%) from responding to users’ requests. Our findings show that while MLLMs can improve visual interpretations’ descriptive accuracy, supporting everyday use also depends on the “visual assistant” skill: behaviors for providing goal-directed, reliable assistance. We conclude by proposing the "visual assistant" skill and guidelines to help MLLM-enabled visual interpretation applications better support BLV people’s access to visual information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7babe47b-6eb8-4fcb-b195-448b7108f196Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable TasksTejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim et al.ICLR 2026 · 154 citations
- Improve Vision Language Model Chain-of-thought ReasoningRuohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang et al.ACL 2025 · 135 citations
- WorldScribe: Towards Context-Aware Live Visual DescriptionsRuei-Che Chang, Yuxuan Liu, Anhong GuoUIST 2024 · 54 citations
- The Emerging Professional Practice of Remote Sighted Assistance for People with Visual ImpairmentsSooyeon Lee, Madison Reddie, Chun-Hua Tsai, Jordan Beck et al.CHI 2020 · 48 citations
Related papers
- Beyond Visual Perception: Insights from Smartphone Interaction of Visually Impaired Users with Large Multimodal ModelsJingyi Xie, Rui Yu, He Zhang, Syed Masum Billah et al.CHI 2025 · 40 citations
- Investigating Use Cases of AI-Powered Scene Description Applications for Blind and Low Vision PeopleRicardo E. Gonzalez Penuela, Jazmin Collins, Cynthia L. Bennett, Shiri AzenkotCHI 2024 · 44 citations
- Understanding the Use of a Large Language Model-Powered Guide to Make Virtual Reality Accessible for Blind and Low Vision PeopleJazmin Collins, Sharon Y. Lin, Tianqi Liu, Andrea Stevenson Won et al.CHI 2026 · 3 citations
- Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired UsersAntonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas et al.ACL 2025
- DR-VQA: Decompose-then-Reconstruct for Visual Question Answering in BLV AssistanceBocheng Pan, Hailong Shi, Xingyu GaoACM MM 2025
