"Rewind to the Jiggling Meat Part": Understanding Voice Control of Instructional Videos in Everyday Tasks
Yaxi Zhao, Razan Jaber, Donald McMillan, Cosmin Munteanu
Abstract
Voice interaction has long been envisioned as enabling users to transform physical interaction into hands-free, such as allowing fine-grained control of instructional videos without physically disengaging from the task at hand. While significant engineering advances have brought us closer to this ideal, we do not fully understand the user requirements for voice interactions that should be supported in such contexts. This paper presents an ecologically-valid wizard-of-oz elicitation study exploring realistic user requirements for an ideal instructional video playback control while cooking. Through the analysis of the issued commands and performed actions during this non-linear and complex task, we identify (1) patterns of command formulation, (2) challenges for design, and (3) how task and voice-based commands are interwoven in real-life. We discuss implications for the design and research of voice interactions for navigating instructional videos while performing complex tasks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e307af84-8f46-4b82-ae77-0d0ccddfbf45Cited by top-tier papers5
- TutoAI: a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasksYuexi Chen, Vlad I. Morariu, Anh Truong, Zhicheng LiuCHI 2024 · 16 citations
- AQuA: Automated Question-Answering in Software Tutorial Videos with Visual AnchorsSaelyne Yang, Jo Vermeulen, George W. Fitzmaurice, Justin MatejkaCHI 2024 · 15 citations
- Stargazer: An Interactive Camera Robot for Capturing How-To Videos Based on Subtle Instructor CuesJiannan Li, Maurício Sousa, Karthik Mahadevan, Bryan Wang et al.CHI 2023 · 12 citations
- Vid2Coach: Transforming How-To Videos into Task AssistantsMina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh et al.UIST 2025 · 9 citations
- "Mango Mango, How to Let The Lettuce Dry Without A Spinner?": Exploring User Perceptions of Using An LLM-Based Conversational Assistant Toward Cooking PartnerSzeyi Chan, Jiachen Li, Bingsheng Yao, Amama Mahmood et al.CSCW 2025 · 4 citations
Related papers
- Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural VideosGeorgianna Lin, Jin Yi Li, Afsaneh Fazly, Vladimir Pavlovic et al.CHI 2023 · 12 citations
- Cooking With Agents: Designing Context-aware Voice InteractionRazan Jaber, Sabrina Zhong, Sanna Kuoppamäki, Aida Hosseini et al.CHI 2024 · 36 citations
- Freehand Grasping: An Analysis of Grasping for Docking Tasks in Virtual RealityAndreea-Dalia Blaga, Maite Frutos-Pascual, Chris Creed, Ian WilliamsIEEE VR 2021 · 18 citations
- Better to Ask Than Assume: Proactive Voice Assistants' Communication Strategies That Respect User Agency in a Smart Home EnvironmentJeesun Oh, Wooseok Kim, Sungbae Kim, Hyeonjeong Im et al.CHI 2024 · 22 citations
- On Pause: How Online Instructional Videos are Used to Achieve Practical TasksSylvaine Tuncer, Barry A. T. Brown, Oskar LindwallCHI 2020 · 30 citations
