(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences
Damien Rudaz, Barbara Nino Carreras, Sara Merlino, Brian L. Due, Barry Brown
摘要
Does human-AI assistance unfold in the same way as humanhuman assistance? This research explores what can be learned from the expertise of blind individuals and sighted volunteers to inform the design of multimodal voice agents and address the enduring challenge of proactivity. Drawing on granular analysis of two representative fragments from a larger corpus, we contrast the practices co-produced by an experienced human remote sighted assistant and a blind participant-as they collaborate to find a stain on a blanket over the phone-with those achieved when the same participant worked with a multimodal voice agent on the same task, a few moments earlier. This comparison enables us to specify precisely which fundamental proactive practices the agent did not enact in situ. We conclude that, so long as multimodal voice agents cannot produce environmentally occasioned vision-based actions, they will lack a key resource relied upon by human remote sighted assistants.
• Human-centered computing → Human computer interaction (HCI); Empirical studies in HCI; Human computer interaction (HCI); Interaction paradigms; Natural language interfaces; Human computer interaction (HCI); HCI theory, concepts and models; Human computer interaction (HCI); HCI design and evaluation methods; Field studies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Human-AI Collaboration via Conditional Delegation: A Case Study of Content ModerationVivian Lai, Samuel Carton, Rajat Bhatnagar, Q. Vera Liao 等CHI 2022 · 被引用 135 次
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu 等CHI 2024 · 被引用 86 次
- The Value, Benefits, and Concerns of Generative AI-Powered Assistance in WritingZhuoyan Li, Chen Liang, Jing Peng, Ming YinCHI 2024 · 被引用 78 次
- Proactive Conversational Agents with Inner ThoughtsXingyu Bruce Liu, Shitao Fang, Weiyan Shi, Chien-Sheng Wu 等CHI 2025 · 被引用 76 次
- The Last Decade of HCI Research on Children and Voice-based Conversational AgentsRadhika Garg, Hua Cui, Spencer Seligson, Bo Zhang 等CHI 2022 · 被引用 71 次
相关 Paper
- The Emerging Professional Practice of Remote Sighted Assistance for People with Visual ImpairmentsSooyeon Lee, Madison Reddie, Chun-Hua Tsai, Jordan Beck 等CHI 2020 · 被引用 48 次
- Support Strategies for Remote Guides in Assisting People with Visual Impairments for Effective Indoor NavigationRie Kamikubo, Naoya Kato, Keita Higuchi, Ryo Yonetani 等CHI 2020 · 被引用 28 次
- Help Supporters: Exploring the Design Space of Assistive Technologies to Support Face-to-Face Help Between Blind and Sighted StrangersYuanyang Teng, Connor Courtien, David Angel Rios, Yves M. Tseng 等CHI 2024 · 被引用 10 次
- Towards More Transactional Voice Assistants: Investigating the Potential for a Multimodal Voice-Activated Indoor Navigation Assistant for Blind and Sighted TravelersAli Abdolrahmani, Maya Howes Gupta, Mei-Lian Vader, Ravi Kuber 等CHI 2021 · 被引用 25 次
- Not Seeing the Whole Picture: Challenges and Opportunities in Using AI for Co-Making Physical, DIY-AT for People with Visual ImpairmentsBen Kosa, Hsuanling Lee, Jasmine Li, Sanbrita Mondal 等CHI 2026 · 被引用 1 次
