(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences
Damien Rudaz, Barbara Nino Carreras, Sara Merlino, Brian L. Due, Barry Brown
Abstract
Does human-AI assistance unfold in the same way as humanhuman assistance? This research explores what can be learned from the expertise of blind individuals and sighted volunteers to inform the design of multimodal voice agents and address the enduring challenge of proactivity. Drawing on granular analysis of two representative fragments from a larger corpus, we contrast the practices co-produced by an experienced human remote sighted assistant and a blind participant-as they collaborate to find a stain on a blanket over the phone-with those achieved when the same participant worked with a multimodal voice agent on the same task, a few moments earlier. This comparison enables us to specify precisely which fundamental proactive practices the agent did not enact in situ. We conclude that, so long as multimodal voice agents cannot produce environmentally occasioned vision-based actions, they will lack a key resource relied upon by human remote sighted assistants.
• Human-centered computing → Human computer interaction (HCI); Empirical studies in HCI; Human computer interaction (HCI); Interaction paradigms; Natural language interfaces; Human computer interaction (HCI); HCI theory, concepts and models; Human computer interaction (HCI); HCI design and evaluation methods; Field studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3169701c-2b06-4907-a592-99041140b575Builds on27
- Human-AI Collaboration via Conditional Delegation: A Case Study of Content ModerationVivian Lai, Samuel Carton, Rajat Bhatnagar, Q. Vera Liao et al.CHI 2022 · 135 citations
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu et al.CHI 2024 · 86 citations
- The Value, Benefits, and Concerns of Generative AI-Powered Assistance in WritingZhuoyan Li, Chen Liang, Jing Peng, Ming YinCHI 2024 · 78 citations
- Proactive Conversational Agents with Inner ThoughtsXingyu Bruce Liu, Shitao Fang, Weiyan Shi, Chien-Sheng Wu et al.CHI 2025 · 76 citations
- The Last Decade of HCI Research on Children and Voice-based Conversational AgentsRadhika Garg, Hua Cui, Spencer Seligson, Bo Zhang et al.CHI 2022 · 71 citations
Related papers
- The Emerging Professional Practice of Remote Sighted Assistance for People with Visual ImpairmentsSooyeon Lee, Madison Reddie, Chun-Hua Tsai, Jordan Beck et al.CHI 2020 · 48 citations
- Support Strategies for Remote Guides in Assisting People with Visual Impairments for Effective Indoor NavigationRie Kamikubo, Naoya Kato, Keita Higuchi, Ryo Yonetani et al.CHI 2020 · 28 citations
- Help Supporters: Exploring the Design Space of Assistive Technologies to Support Face-to-Face Help Between Blind and Sighted StrangersYuanyang Teng, Connor Courtien, David Angel Rios, Yves M. Tseng et al.CHI 2024 · 10 citations
- Towards More Transactional Voice Assistants: Investigating the Potential for a Multimodal Voice-Activated Indoor Navigation Assistant for Blind and Sighted TravelersAli Abdolrahmani, Maya Howes Gupta, Mei-Lian Vader, Ravi Kuber et al.CHI 2021 · 25 citations
- Not Seeing the Whole Picture: Challenges and Opportunities in Using AI for Co-Making Physical, DIY-AT for People with Visual ImpairmentsBen Kosa, Hsuanling Lee, Jasmine Li, Sanbrita Mondal et al.CHI 2026 · 1 citation
