AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multimodal Information Between Reality and Videos
Zheng Ning, Leyang Li, Daniel Killough, JooYoung Seo, Patrick Carrington, Yapeng Tian, Yuhang Zhao, Franklin Mingzhe Li, Toby Jia-Jun Li
摘要
Videos offer rich audiovisual information that can support people in performing activities of daily living (ADLs), but they remain largely inaccessible to blind or low-vision (BLV) individuals.In cooking, BLV people often rely on non-visual cues-such as touch, taste, and smell-to navigate their environment, making it difficult to follow UIST '25, September 28-October 01, 2025, Busan, Republic of Korea Ning et al.the predominantly audiovisual instructions found in video recipes.To address this problem, we introduce Aroma, an AI system that provides timely responses to the user based on real-time, contextaware assistance by integrating non-visual cues perceived by the user, a wearable camera feed, and video recipe content.Aroma uses a mixed-initiative approach: it responds to user requests while also proactively monitoring the video stream to offer timely alerts and guidance.This collaborative design leverages the complementary strengths of the user and AI system to align the physical environment with the video recipe, helping the user interpret their current state and make sense of the steps.We evaluated Aroma through a study with eight BLV participants and offered insights for designing interactive AI systems to support BLV individuals in performing ADLs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ADCanvas: Accessible and Conversational Audio Description Authoring for Blind and Low Vision CreatorsFranklin Mingzhe Li, Michael Xieyang Liu, Cynthia L. Bennett, Shaun K. KaneCHI 2026 · 被引用 2 次
- Co-Designing Multimodal Systems for Accessible Asynchronous Dance InstructionUjjaini Das, Shreya Kappala, Meng Chen, Mina Huh 等CHI 2026 · 被引用 1 次
- Lost in Instructions: Study of Blind Users' Experiences with DIY Manuals and AI-Rewritten Instructions for Assembly, Operation, and Troubleshooting of Tangible ProductsMonalika Padma Reddy, Aruna Balasubramanian, Jiawei Zhou, Xiaojun Bi 等CHI 2026 · 被引用 1 次
- Understanding Nature Engagement Experiences of Blind PeopleMengjie Tang, Xinman Li, Juxiao Zhang, Franklin Mingzhe Li 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper22
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等ICCV 2021 · 被引用 345 次
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu 等CVPR 2022 · 被引用 101 次
- Toward Automatic Audio Description Generation for Accessible VideosYujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang 等CHI 2021 · 被引用 86 次
- What Makes Videos Accessible to Blind and Visually Impaired People?Xingyu Liu, Patrick Carrington, Xiang 'Anthony' Chen, Amy PavelCHI 2021 · 被引用 78 次
- EarVR: Using Ear Haptics in Virtual Reality for Deaf and Hard-of-Hearing PeopleMohammadreza Mirzaei, Peter Kán, Hannes KaufmannIEEE VR 2020 · 被引用 73 次
相关 Paper
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh 等UIST 2025 · 被引用 9 次
- CookAR: Affordance Augmentations in Wearable AR to Support Kitchen Tool Interactions for People with Low VisionJaewook Lee, Andrew D. Tjahjadi, Jiho Kim, Junpu Yu 等UIST 2024 · 被引用 25 次
- Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural VideosGeorgianna Lin, Jin Yi Li, Afsaneh Fazly, Vladimir Pavlovic 等CHI 2023 · 被引用 12 次
- "It's Kind of Context Dependent": Understanding Blind and Low Vision People's Video Accessibility Preferences Across Viewing ScenariosLucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl 等CHI 2024 · 被引用 23 次
- SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersZheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen 等CHI 2024 · 被引用 27 次
