TutoAI: a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasks
Yuexi Chen, Vlad I. Morariu, Anh Truong, Zhicheng Liu
Abstract
Mixed-media tutorials, which integrate videos, images, text, and diagrams to teach procedural skills, offer more browsable alternatives than timeline-based videos. However, manually creating such tutorials is tedious, and existing automated solutions are often restricted to a particular domain. While AI models hold promise, it is unclear how to effectively harness their powers, given the multi-modal data involved and the vast landscape of models. We present TutoAI, a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasks. First, we distill common tutorial components by surveying existing work; then, we present an approach to identify, assemble, and evaluate AI models for component extraction; finally, we propose guidelines for designing user interfaces (UI) that support tutorial creation based on AI-generated components. We show that TutoAI has achieved higher or similar quality compared to a baseline model in preliminary user studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 609f5d44-ad16-4e57-b659-3f5fb82f8bcdCited by top-tier papers4
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 33 citations
- NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video UnderstandingRunning Zhao, Zhihan Jiang, Xinchen Zhang, Chirui Chang et al.UIST 2025 · 6 citations
- 'How Do You Know That Stuff?': Barriers to Expertise Sharing Among Spreadsheet UsersQing (Nancy) Xia, Advait Sarkar, Duncan P. Brumby, Anna L. CoxCSCW 2025 · 2 citations
- ClearFairy: Capturing Creative Workflows through Decision Structuring, In-Situ Questioning, and Rationale InferenceKihoon Son, DaEun Choi, Tae Soo Kim, Young-Ho Kim et al.CHI 2026 · 1 citation
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- Automatic Instructional Video Creation from a Markdown-Formatted TutorialPeggy Chi, Nathan Frey, Katrina Panovich, Irfan EssaUIST 2021 · 27 citations
- Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup VideosAnh Truong, Peggy Chi, David Salesin, Irfan Essa et al.CHI 2021 · 57 citations
- ARify: Leveraging Narrated Instructional Videos to Create Augmented Reality Tutorials for Procedural TasksXiyun Hu, Chenfei Zhu, Shao-Kang Hsia, Dizhi Ma et al.CHI 2026 · 1 citation
- Demo2Tutorial: From Human Experience to Multimodal Software TutorialsZechen Bai, Zhiheng Chen, Yiqi Lin, Kevin Qinghong Lin et al.CVPR 2026
- Screencast Tutorial Video UnderstandingKunpeng Li, Chen Fang, Zhaowen Wang, Seokhwan Kim et al.CVPR 2020
