RubySlippers: Supporting Content-based Voice Navigation for How-to Videos
Minsuk Chang, Mina Huh, Juho Kim
Abstract
Directly manipulating the timeline, such as scrubbing for thumbnails, is the standard way of controlling how-to videos. However, when how-to videos involve physical activities, people inconveniently alternate between controlling the video and performing the tasks. Adopting a voice user interface allows people to control the video with voice while performing the tasks with hands. However, naively translating timeline manipulation into voice user interfaces (VUI) results in temporal referencing (e.g. “rewind 20 seconds”), which requires a different mental model for navigation and thereby limiting users’ ability to peek into the content. We present RubySlippers, a system that supports efficient content-based voice navigation through keyword-based queries. Our computational pipeline automatically detects referenceable elements in the video, and finds the video segmentation that minimizes the number of needed navigational commands. Our evaluation (N=12) shows that participants could perform three representative navigation tasks with fewer commands and less frustration using RubySlippers than the conventional voice-enabled video interface.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bea33fff-7b62-4cf9-ae89-bb4dd5b82440Cited by top-tier papers12
- A Literature Review of Video-Sharing Platform Research in HCIAva Bartolome, Shuo NiuCHI 2023 · 63 citations
- AVscript: Accessible Video Editing with Audio-Visual ScriptsMina Huh, Saelyne Yang, Yi-Hao Peng, Xiang 'Anthony' Chen et al.CHI 2023 · 44 citations
- "It Feels Like Taking a Gamble": Exploring Perceptions, Practices, and Challenges of Using Makeup and Cosmetics for People with Visual ImpairmentsFranklin Mingzhe Li, Franchesca Spektor, Meng Xia, Mina Huh et al.CHI 2022 · 39 citations
- VideoDiff: Human-AI Video Co-Creation with AlternativesMina Huh, Ding Li, Kim Pimmel, Hijung Valentina Shin et al.CHI 2025 · 26 citations
- Synthesis-Assisted Video Prototyping From a DocumentPeggy Chi, Tao Dong, Christian Früh, Brian Colonna et al.UIST 2022 · 18 citations
Related papers
- Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup VideosAnh Truong, Peggy Chi, David Salesin, Irfan Essa et al.CHI 2021 · 57 citations
- "Rewind to the Jiggling Meat Part": Understanding Voice Control of Instructional Videos in Everyday TasksYaxi Zhao, Razan Jaber, Donald McMillan, Cosmin MunteanuCHI 2022 · 18 citations
- Identifying Multimodal Context Awareness Requirements for Supporting User Interaction with Procedural VideosGeorgianna Lin, Jin Yi Li, Afsaneh Fazly, Vladimir Pavlovic et al.CHI 2023 · 12 citations
- Reactive Video: Adaptive Video Playback Based on User Motion for Supporting Physical ActivityChristopher Clarke, Doga Cavdir, Patrick Chiu, Laurent Denoue et al.UIST 2020 · 30 citations
- Automatic Instructional Video Creation from a Markdown-Formatted TutorialPeggy Chi, Nathan Frey, Katrina Panovich, Irfan EssaUIST 2021 · 27 citations
