Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural Videos
Georgianna Lin, Jin Yi Li, Afsaneh Fazly, Vladimir Pavlovic, Khai N. Truong
Abstract
Following along how-to videos requires alternating focus between understanding procedural video instructions and performing them. Examining how to support these continuous context switches for the user has been largely unexplored. In this paper, we describe a user study with thirty participants who performed an hour-long cooking task while interacting with a wizard-of-oz hands-free interactive system that is aware of both their cooking progress and environment contexts. Through analysis of the session scripts, we identify a dichotomy between participant query differences and workflow alignment similarities, under-studied interactions that require AI functionality beyond video navigation alone, and queries that call for multimodal sensing of a user’s environment. By understanding the assistant experience through the participants’ interactions, we identify design implications for a smart assistant that can discern a user’s task completion flow and personal characteristics, accommodate requests within and external to the task domain, and support nonvoice-based queries.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b8f10996-7192-4d4e-98a4-fa42a1bc68c5Cited by top-tier papers3
- PrISM-Q&A: Step-Aware Voice Assistant on a Smartwatch Enabled by Multimodal Procedure Tracking and Large Language ModelsRiku Arakawa, Jill Fain Lehman, Mayank GoelUbiComp 2025 · 20 citations
- AQuA: Automated Question-Answering in Software Tutorial Videos with Visual AnchorsSaelyne Yang, Jo Vermeulen, George W. Fitzmaurice, Justin MatejkaCHI 2024 · 15 citations
- AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multimodal Information Between Reality and VideosZheng Ning, Leyang Li, Daniel Killough, JooYoung Seo et al.UIST 2025 · 1 citation
Related papers
- "Rewind to the Jiggling Meat Part": Understanding Voice Control of Instructional Videos in Everyday TasksYaxi Zhao, Razan Jaber, Donald McMillan, Cosmin MunteanuCHI 2022 · 18 citations
- Cooking With Agents: Designing Context-aware Voice InteractionRazan Jaber, Sabrina Zhong, Sanna Kuoppamäki, Aida Hosseini et al.CHI 2024 · 36 citations
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh et al.UIST 2025 · 9 citations
- AMMA: Adaptive Multimodal Assistants Through Automated State Tracking and User Model-Directed Guidance PlanningJackie (Junrui) Yang, Leping Qiu, Emmanuel Angel Corona-Moreno, Louisa Shi et al.IEEE VR 2024 · 11 citations
- Better to Ask Than Assume: Proactive Voice Assistants' Communication Strategies That Respect User Agency in a Smart Home EnvironmentJeesun Oh, Wooseok Kim, Sungbae Kim, Hyeonjeong Im et al.CHI 2024 · 22 citations
