Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural Videos
Georgianna Lin, Jin Yi Li, Afsaneh Fazly, Vladimir Pavlovic, Khai N. Truong
摘要
Following along how-to videos requires alternating focus between understanding procedural video instructions and performing them. Examining how to support these continuous context switches for the user has been largely unexplored. In this paper, we describe a user study with thirty participants who performed an hour-long cooking task while interacting with a wizard-of-oz hands-free interactive system that is aware of both their cooking progress and environment contexts. Through analysis of the session scripts, we identify a dichotomy between participant query differences and workflow alignment similarities, under-studied interactions that require AI functionality beyond video navigation alone, and queries that call for multimodal sensing of a user’s environment. By understanding the assistant experience through the participants’ interactions, we identify design implications for a smart assistant that can discern a user’s task completion flow and personal characteristics, accommodate requests within and external to the task domain, and support nonvoice-based queries.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- PrISM-Q&A: Step-Aware Voice Assistant on a Smartwatch Enabled by Multimodal Procedure Tracking and Large Language ModelsRiku Arakawa, Jill Fain Lehman, Mayank GoelUbiComp 2025 · 被引用 20 次
- AQuA: Automated Question-Answering in Software Tutorial Videos with Visual AnchorsSaelyne Yang, Jo Vermeulen, George W. Fitzmaurice, Justin MatejkaCHI 2024 · 被引用 15 次
- AROMA: Mixed-Initiative AI Assistance for Non-Visual Cooking by Grounding Multimodal Information Between Reality and VideosZheng Ning, Leyang Li, Daniel Killough, JooYoung Seo 等UIST 2025 · 被引用 1 次
相关 Paper
- "Rewind to the Jiggling Meat Part": Understanding Voice Control of Instructional Videos in Everyday TasksYaxi Zhao, Razan Jaber, Donald McMillan, Cosmin MunteanuCHI 2022 · 被引用 18 次
- Cooking With Agents: Designing Context-aware Voice InteractionRazan Jaber, Sabrina Zhong, Sanna Kuoppamäki, Aida Hosseini 等CHI 2024 · 被引用 36 次
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh 等UIST 2025 · 被引用 9 次
- AMMA: Adaptive Multimodal Assistants Through Automated State Tracking and User Model-Directed Guidance PlanningJackie (Junrui) Yang, Leping Qiu, Emmanuel Angel Corona-Moreno, Louisa Shi 等IEEE VR 2024 · 被引用 11 次
- Better to Ask Than Assume: Proactive Voice Assistants' Communication Strategies That Respect User Agency in a Smart Home EnvironmentJeesun Oh, Wooseok Kim, Sungbae Kim, Hyeonjeong Im 等CHI 2024 · 被引用 22 次
