HelpViz: Automatic Generation of Contextual Visual Mobile Tutorials from Text-Based Instructions
Mingyuan Zhong, Gang Li, Peggy Chi, Yang Li
Abstract
We present HelpViz, a tool for generating contextual visual mobile tutorials from text-based instructions that are abundant on the web. HelpViz transforms text instructions to graphical tutorials in batch, by extracting a sequence of actions from each text instruction through an instruction parsing model, and executing the extracted actions on a simulation infrastructure that manages an array of Android emulators. The automatic execution of each instruction produces a set of graphical and structural assets, including images, videos, and metadata such as clicked elements for each step. HelpViz then synthesizes a tutorial by combining parsed text instructions with the generated assets, and contextualizes the tutorial to user interaction by tracking the user's progress and highlighting the next step. Our experiments with HelpViz indicate that our pipeline improved tutorial execution robustness and that participants preferred tutorials generated by HelpViz over text-based instructions. HelpViz promises a cost-effective approach for generating contextual visual tutorials for mobile interaction at scale.
• Human-centered computing → Interactive systems and tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 17b6d18e-609c-45c2-aa9b-1bbbc4b5d66eCited by top-tier papers6
- HelpCall: Designing Informal Technology Assistance for Older Adults via VideoconferencingTeerapaun Tanprasert, Jiamin Dai, Joanna McGrenereCHI 2024 · 20 citations
- Synthesis-Assisted Video Prototyping From a DocumentPeggy Chi, Tao Dong, Christian Früh, Brian Colonna et al.UIST 2022 · 18 citations
- ScreenAudit: Detecting Screen Reader Accessibility Errors in Mobile Apps Using Large Language ModelsMingyuan Zhong, Ruolin Chen, Xia Chen, James Fogarty et al.CHI 2025 · 15 citations
- Video2Action: Reducing Human Interactions in Action Annotation of App Tutorial VideosSidong Feng, Chunyang Chen, Zhenchang XingUIST 2023 · 12 citations
- DuetUI: A Bidirectional Context Loop for Human-Agent Co-Generation of Task-Oriented InterfacesYuan Xu, Shaowen Xiang, Yizhi Song, Ruoting Sun et al.CHI 2026 · 2 citations
Builds on5
- Mapping Natural Language Instructions to Mobile UI Action SequencesYang Li, Jiacong He, Xin Zhou, Yuan Zhang et al.ACL 2020 · 75 citations
- Automatic Generation of Two-Level Hierarchical Tutorials from Instructional Makeup VideosAnh Truong, Peggy Chi, David Salesin, Irfan Essa et al.CHI 2021 · 57 citations
- Generating Audio-Visual Slideshows from Text Articles Using Word ConcretenessMackenzie Leake, Hijung Valentina Shin, Joy O. Kim, Maneesh AgrawalaCHI 2020 · 36 citations
- Composing Flexibly-Organized Step-by-Step Tutorials from Linked Source Code, Snippets, and OutputsAndrew Head, Jason Jiang, James Smith, Marti A. Hearst et al.CHI 2020 · 29 citations
- If beam search is the answer, what was the question?Clara Meister, Ryan Cotterell, Tim VieiraEMNLP 2020 · 26 citations
Related papers
- Automatic Instructional Video Creation from a Markdown-Formatted TutorialPeggy Chi, Nathan Frey, Katrina Panovich, Irfan EssaUIST 2021 · 27 citations
- TutoAI: a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasksYuexi Chen, Vlad I. Morariu, Anh Truong, Zhicheng LiuCHI 2024 · 16 citations
- Harnessing Webpage UIs for Text-Rich Visual UnderstandingJunpeng Liu, Tianyue Ou, Yifan Song, Yuxiao Qu et al.ICLR 2025
- AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web TutorialsYiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang et al.ICLR 2025
- Stitch-a-Demo: Creating Video Demonstrations from Multistep DescriptionsChi Hsuan Wu, Kumar Ashutosh, Kristen GraumanCVPR 2026 · 1 citation
