Beyond Visual Perception: Insights from Smartphone Interaction of Visually Impaired Users with Large Multimodal Models
Jingyi Xie, Rui Yu, He Zhang, Syed Masum Billah, Sooyeon Lee, John M. Carroll
Abstract
Large multimodal models (LMMs) have enabled new AI-powered applications that help people with visual impairments (PVI) receive natural language descriptions of their surroundings through audible text. We investigated how this emerging paradigm of visual assistance transforms how PVI perform and manage their daily tasks. Moving beyond usability assessments, we examined both the capabilities and limitations of LMM-based tools in personal and social contexts, while exploring design implications for their future development. Through interviews with 14 visually impaired users of Be My AI (an LMM-based application) and analysis of its image descriptions from both study participants and social media platforms, we identified two key limitations. First, these systems' context awareness suffers from hallucinations and misinterpretations of social contexts, styles, and human identities. Second, their intent-oriented capabilities often fail to grasp and act on users' intentions. Based on these findings, we propose design strategies for improving both human-AI and AI-AI interactions, contributing to the development of more effective, interactive, and personalized assistive technologies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b1f707c-fb1c-4b89-86c6-7babd7003e3bCited by top-tier papers6
- (Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection SequencesDamien Rudaz, Barbara Nino Carreras, Sara Merlino, Brian L. Due et al.CHI 2026 · 2 citations
- TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual DescriptionsRuei-Che Chang, Rosiana Natalie, Wenqian Xu, Jovan Zheng Feng Yap et al.CHI 2026 · 2 citations
- "It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language ModelsKapil Garg, Xinru Tang, Jimin Heo, Dwayne R. Morgan et al.CHI 2026 · 2 citations
- How Multimodal Large Language Models Support Access to Visual Information: A Diary Study With Blind and Low Vision PeopleRicardo E. Gonzalez Penuela, Crescentia Jung, Sharon Y. Lin, Ruiying Hu et al.CHI 2026 · 1 citation
- Lost in Instructions: Study of Blind Users' Experiences with DIY Manuals and AI-Rewritten Instructions for Assembly, Operation, and Troubleshooting of Tangible ProductsMonalika Padma Reddy, Aruna Balasubramanian, Jiawei Zhou, Xiaojun Bi et al.CHI 2026 · 1 citation
Builds on20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le et al.ICLR 2023 · 681 citations
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
Related papers
- Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired UsersAntonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas et al.ACL 2025
- Understanding the Use of a Large Language Model-Powered Guide to Make Virtual Reality Accessible for Blind and Low Vision PeopleJazmin Collins, Sharon Y. Lin, Tianqi Liu, Andrea Stevenson Won et al.CHI 2026 · 3 citations
- "This is My Fault", Really? Understanding Blind and Low-Vision People's Perception of Hallucination in Large Vision Language ModelsYilin Tang, Yuyang Fang, Tianle Wang, Lingyun Sun et al.UIST 2025 · 3 citations
- Not Seeing the Whole Picture: Challenges and Opportunities in Using AI for Co-Making Physical, DIY-AT for People with Visual ImpairmentsBen Kosa, Hsuanling Lee, Jasmine Li, Sanbrita Mondal et al.CHI 2026 · 1 citation
- Toward Independent Online Shopping of the Visually Impaired Through Voice-based Computer-Using AgentSubin Shin, Jeesun Oh, Suhyun Kim, Seoyeon Eom et al.CHI 2026 · 1 citation
