Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup Videos
Anh Truong, Peggy Chi, David Salesin, Irfan Essa, Maneesh Agrawala
Abstract
We present a multi-modal approach for automatically generating hierarchical tutorials from instructional makeup videos. Our approach is inspired by prior research in cognitive psychology, which suggests that people mentally segment procedural tasks into event hierarchies, where coarse-grained events focus on objects while fine-grained events focus on actions. In the instructional makeup domain, we find that objects correspond to facial parts while fine-grained steps correspond to actions on those facial parts. Given an input instructional makeup video, we apply a set of heuristics that combine computer vision techniques with transcript text analysis to automatically identify the fine-level action steps and group these steps by facial part to form the coarse-level events. We provide a voice-enabled, mixed-media UI to visualize the resulting hierarchy and allow users to efficiently navigate the tutorial (e.g., skip ahead, return to previous steps) at their own pace. Users can navigate the hierarchy at both the facial-part and action-step levels using click-based interactions and voice commands. We demonstrate the effectiveness of segmentation algorithms and the resulting mixed-media UI on a variety of input makeup videos. A user study shows that users prefer following instructional makeup videos in our mixed-media format to the standard video UI and that they find our format much easier to navigate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d0b64b8-f217-462c-908e-8a8bcd275d30Cited by top-tier papers27
- A Literature Review of Video-Sharing Platform Research in HCIAva Bartolome, Shuo NiuCHI 2023 · 63 citations
- AVscript: Accessible Video Editing with Audio-Visual ScriptsMina Huh, Saelyne Yang, Yi-Hao Peng, Xiang 'Anthony' Chen et al.CHI 2023 · 44 citations
- "It Feels Like Taking a Gamble": Exploring Perceptions, Practices, and Challenges of Using Makeup and Cosmetics for People with Visual ImpairmentsFranklin Mingzhe Li, Franchesca Spektor, Meng Xia, Mina Huh et al.CHI 2022 · 39 citations
- HelpViz: Automatic Generation of Contextual Visual Mobile Tutorials from Text-Based InstructionsMingyuan Zhong, Gang Li, Peggy Chi, Yang LiUIST 2021 · 27 citations
- A Layered Authoring Tool for Stylized 3D animationsJiaju Ma, Li-Yi Wei, Rubaiat Habib KaziCHI 2022 · 22 citations
Builds on4
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Temporal Segmentation of Creative Live StreamsC. Ailie Fraser, Joy O. Kim, Hijung Valentina Shin, Joel Brandt et al.CHI 2020 · 38 citations
- On Pause: How Online Instructional Videos are Used to Achieve Practical TasksSylvaine Tuncer, Barry A. T. Brown, Oskar LindwallCHI 2020 · 30 citations
- Instructional Video Design: Investigating the Impact of Monologue- and Dialogue-style PresentationsBridjet Lee, Kasia MuldnerCHI 2020 · 24 citations
Related papers
- Automatic Instructional Video Creation from a Markdown-Formatted TutorialPeggy Chi, Nathan Frey, Katrina Panovich, Irfan EssaUIST 2021 · 27 citations
- TutoAI: a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasksYuexi Chen, Vlad I. Morariu, Anh Truong, Zhicheng LiuCHI 2024 · 16 citations
- Screencast Tutorial Video UnderstandingKunpeng Li, Chen Fang, Zhaowen Wang, Seokhwan Kim et al.CVPR 2020
- SeeHow: Workflow Extraction from Programming Screencasts through Action-Aware Video AnalyticsDehai Zhao, Zhenchang Xing, Xin Xia, Deheng Ye et al.ICSE 2023 · 6 citations
- Learning To Recognize Procedural Activities with Distant SupervisionXudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach et al.CVPR 2022 · 55 citations
