Synthesis-Assisted Video Prototyping From a Document
Peggy Chi, Tao Dong, Christian Früh, Brian Colonna, Vivek Kwatra, Irfan Essa
Abstract
Video productions commonly start with a script, especially for talking head videos that feature a speaker narrating to the camera. When the source materials come from a written document – such as a web tutorial, it takes iterations to refine content from a text article to a spoken dialogue, while considering visual compositions in each scene. We propose Doc2Video, a video prototyping approach that converts a document to interactive scripting with a preview of synthetic talking head videos. Our pipeline decomposes a source document into a series of scenes, each automatically creating a synthesized video of a virtual instructor. Designed for a specific domain – programming cookbooks, we apply visual elements from the source document, such as a keyword, a code snippet or a screenshot, in suitable layouts. Users edit narration sentences, break or combine sections, and modify visuals to prototype a video in our Editing UI. We evaluated our pipeline with public programming cookbooks. Feedback from professional creators shows that our method provided a reasonable starting point to engage them in interactive scripting for a narrated instructional video.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5116f134-c894-4b64-8bd2-cb3995b4beb6Cited by top-tier papers8
- DataParticles: Block-based and Language-oriented Authoring of Animated Unit VisualizationsYining Cao, Jane L. E, Chen Zhu-Tian, Haijun XiaCHI 2023 · 53 citations
- VideoDiff: Human-AI Video Co-Creation with AlternativesMina Huh, Ding Li, Kim Pimmel, Hijung Valentina Shin et al.CHI 2025 · 26 citations
- Papeos: Augmenting Research Papers with Talk VideosTae Soo Kim, Matt Latzke, Jonathan Bragg, Amy X. Zhang et al.UIST 2023 · 17 citations
- Compositional Structures as Substrates for Human-AI Co-creation Environment: A Design Approach and A Case StudyYining Cao, Yiyi Huang, Anh Truong, Hijung Valentina Shin et al.CHI 2025 · 15 citations
- Reflecting on Design Paradigms of Animated Data Video ToolsLeixian Shen, Haotian Li, Yun Wang, Huamin QuCHI 2025 · 10 citations
Builds on14
- Rescribe: Authoring and Automatically Editing Audio DescriptionsAmy Pavel, Gabriel Reyes, Jeffrey P. BighamUIST 2020 · 72 citations
- Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup VideosAnh Truong, Peggy Chi, David Salesin, Irfan Essa et al.CHI 2021 · 57 citations
- Towards Supporting Programming Education at Scale via Live StreamingYan Chen, Walter S. Lasecki, Tao DongCSCW 2020 · 45 citations
- Crosscast: Adding Visuals to Audio Travel PodcastsHaijun Xia, Jennifer Jacobs, Maneesh AgrawalaUIST 2020 · 44 citations
- RubySlippers: Supporting Content-based Voice Navigation for How-to VideosMinsuk Chang, Mina Huh, Juho KimCHI 2021 · 40 citations
Related papers
- Automatic Instructional Video Creation from a Markdown-Formatted TutorialPeggy Chi, Nathan Frey, Katrina Panovich, Irfan EssaUIST 2021 · 27 citations
- Code2Video: A Code-centric Paradigm for Educational Video CreationYanzhe Chen, Kevin Qinghong Lin, Mike Zheng ShouICML 2026
- Automatic Video Creation From a Web PagePeggy Chi, Zheng Sun, Katrina Panovich, Irfan EssaUIST 2020 · 25 citations
- PaperTok: Exploring the Use of Generative AI for Creating Short-form Videos for Research CommunicationMeziah Ruby Cristobal, Hyeonjeong Byeon, Tze-Yu Chen, Ruoxi Shang et al.CHI 2026 · 2 citations
- Demo2Tutorial: From Human Experience to Multimodal Software TutorialsZechen Bai, Zhiheng Chen, Yiqi Lin, Kevin Qinghong Lin et al.CVPR 2026
