AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multimodal Information Between Reality and Videos
Zheng Ning, Leyang Li, Daniel Killough, JooYoung Seo, Patrick Carrington, Yapeng Tian, Yuhang Zhao, Franklin Mingzhe Li, Toby Jia-Jun Li
Abstract
Videos offer rich audiovisual information that can support people in performing activities of daily living (ADLs), but they remain largely inaccessible to blind or low-vision (BLV) individuals.In cooking, BLV people often rely on non-visual cues-such as touch, taste, and smell-to navigate their environment, making it difficult to follow UIST '25, September 28-October 01, 2025, Busan, Republic of Korea Ning et al.the predominantly audiovisual instructions found in video recipes.To address this problem, we introduce Aroma, an AI system that provides timely responses to the user based on real-time, contextaware assistance by integrating non-visual cues perceived by the user, a wearable camera feed, and video recipe content.Aroma uses a mixed-initiative approach: it responds to user requests while also proactively monitoring the video stream to offer timely alerts and guidance.This collaborative design leverages the complementary strengths of the user and AI system to align the physical environment with the video recipe, helping the user interpret their current state and make sense of the steps.We evaluated Aroma through a study with eight BLV participants and offered insights for designing interactive AI systems to support BLV individuals in performing ADLs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a77dba58-7ba6-4f2e-a98d-306ea54bfe0aCited by top-tier papers4
- ADCanvas: Accessible and Conversational Audio Description Authoring for Blind and Low Vision CreatorsFranklin Mingzhe Li, Michael Xieyang Liu, Cynthia L. Bennett, Shaun K. KaneCHI 2026 · 2 citations
- Co-Designing Multimodal Systems for Accessible Asynchronous Dance InstructionUjjaini Das, Shreya Kappala, Meng Chen, Mina Huh et al.CHI 2026 · 1 citation
- Lost in Instructions: Study of Blind Users' Experiences with DIY Manuals and AI-Rewritten Instructions for Assembly, Operation, and Troubleshooting of Tangible ProductsMonalika Padma Reddy, Aruna Balasubramanian, Jiawei Zhou, Xiaojun Bi et al.CHI 2026 · 1 citation
- Understanding Nature Engagement Experiences of Blind PeopleMengjie Tang, Xinman Li, Juxiao Zhang, Franklin Mingzhe Li et al.CHI 2026 · 1 citation
Builds on22
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.ICCV 2021 · 345 citations
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
- Toward Automatic Audio Description Generation for Accessible VideosYujia Wang, Wei Liang, Haikun Huang, Yongqi Zhang et al.CHI 2021 · 86 citations
- What Makes Videos Accessible to Blind and Visually Impaired People?Xingyu Liu, Patrick Carrington, Xiang 'Anthony' Chen, Amy PavelCHI 2021 · 78 citations
- EarVR: Using Ear Haptics in Virtual Reality for Deaf and Hard-of-Hearing PeopleMohammadreza Mirzaei, Peter Kán, Hannes KaufmannIEEE VR 2020 · 73 citations
Related papers
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh et al.UIST 2025 · 9 citations
- CookAR: Affordance Augmentations in Wearable AR to Support Kitchen Tool Interactions for People with Low VisionJaewook Lee, Andrew D. Tjahjadi, Jiho Kim, Junpu Yu et al.UIST 2024 · 25 citations
- Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural VideosGeorgianna Lin, Jin Yi Li, Afsaneh Fazly, Vladimir Pavlovic et al.CHI 2023 · 12 citations
- "It's Kind of Context Dependent": Understanding Blind and Low Vision People's Video Accessibility Preferences Across Viewing ScenariosLucy Jiang, Crescentia Jung, Mahika Phutane, Abigale Stangl et al.CHI 2024 · 23 citations
- SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision ViewersZheng Ning, Brianna L. Wimer, Kaiwen Jiang, Keyi Chen et al.CHI 2024 · 27 citations
