Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup Videos
Anh Truong, Peggy Chi, David Salesin, Irfan Essa, Maneesh Agrawala
摘要
We present a multi-modal approach for automatically generating hierarchical tutorials from instructional makeup videos. Our approach is inspired by prior research in cognitive psychology, which suggests that people mentally segment procedural tasks into event hierarchies, where coarse-grained events focus on objects while fine-grained events focus on actions. In the instructional makeup domain, we find that objects correspond to facial parts while fine-grained steps correspond to actions on those facial parts. Given an input instructional makeup video, we apply a set of heuristics that combine computer vision techniques with transcript text analysis to automatically identify the fine-level action steps and group these steps by facial part to form the coarse-level events. We provide a voice-enabled, mixed-media UI to visualize the resulting hierarchy and allow users to efficiently navigate the tutorial (e.g., skip ahead, return to previous steps) at their own pace. Users can navigate the hierarchy at both the facial-part and action-step levels using click-based interactions and voice commands. We demonstrate the effectiveness of segmentation algorithms and the resulting mixed-media UI on a variety of input makeup videos. A user study shows that users prefer following instructional makeup videos in our mixed-media format to the standard video UI and that they find our format much easier to navigate.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- A Literature Review of Video-Sharing Platform Research in HCIAva Bartolome, Shuo NiuCHI 2023 · 被引用 63 次
- AVscript: Accessible Video Editing with Audio-Visual ScriptsMina Huh, Saelyne Yang, Yi-Hao Peng, Xiang 'Anthony' Chen 等CHI 2023 · 被引用 44 次
- "It Feels Like Taking a Gamble": Exploring Perceptions, Practices, and Challenges of Using Makeup and Cosmetics for People with Visual ImpairmentsFranklin Mingzhe Li, Franchesca Spektor, Meng Xia, Mina Huh 等CHI 2022 · 被引用 39 次
- HelpViz: Automatic Generation of Contextual Visual Mobile Tutorials from Text-Based InstructionsMingyuan Zhong, Gang Li, Peggy Chi, Yang LiUIST 2021 · 被引用 27 次
- A Layered Authoring Tool for Stylized 3D animationsJiaju Ma, Li-Yi Wei, Rubaiat Habib KaziCHI 2022 · 被引用 22 次
它引用的顶会 Paper4
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Temporal Segmentation of Creative Live StreamsC. Ailie Fraser, Joy O. Kim, Hijung Valentina Shin, Joel Brandt 等CHI 2020 · 被引用 38 次
- On Pause: How Online Instructional Videos are Used to Achieve Practical TasksSylvaine Tuncer, Barry A. T. Brown, Oskar LindwallCHI 2020 · 被引用 30 次
- Instructional Video Design: Investigating the Impact of Monologue- and Dialogue-style PresentationsBridjet Lee, Kasia MuldnerCHI 2020 · 被引用 24 次
相关 Paper
- Automatic Instructional Video Creation from a Markdown-Formatted TutorialPeggy Chi, Nathan Frey, Katrina Panovich, Irfan EssaUIST 2021 · 被引用 27 次
- TutoAI: a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasksYuexi Chen, Vlad I. Morariu, Anh Truong, Zhicheng LiuCHI 2024 · 被引用 16 次
- Screencast Tutorial Video UnderstandingKunpeng Li, Chen Fang, Zhaowen Wang, Seokhwan Kim 等CVPR 2020
- SeeHow: Workflow Extraction from Programming Screencasts through Action-Aware Video AnalyticsDehai Zhao, Zhenchang Xing, Xin Xia, Deheng Ye 等ICSE 2023 · 被引用 6 次
- Learning To Recognize Procedural Activities with Distant SupervisionXudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach 等CVPR 2022 · 被引用 55 次
