AQuA: Automated Question-Answering in Software Tutorial Videos with Visual Anchors
Saelyne Yang, Jo Vermeulen, George W. Fitzmaurice, Justin Matejka
Abstract
Tutorial videos are a popular help source for learning feature-rich software. However, getting quick answers to questions about tutorial videos is difficult. We present an automated approach for responding to tutorial questions. By analyzing 633 questions found in 5,944 video comments, we identified different question types and observed that users frequently described parts of the video in questions. We then asked participants (N=24) to watch tutorial videos and ask questions while annotating the video with relevant visual anchors. Most visual anchors referred to UI elements and the application workspace. Based on these insights, we built AQuA, a pipeline that generates useful answers to questions with visual anchors. We demonstrate this for Fusion 360, showing that we can recognize UI elements in visual anchors and generate answers using GPT-4 augmented with that visual information and software documentation. An evaluation study (N=16) demonstrates that our approach provides better answers than baseline methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd34d2c2-0467-4be4-9fd0-40fa0c934001Cited by top-tier papers4
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 33 citations
- Social-RAG: Retrieving from Group Interactions to Socially Ground AI GenerationRuotong Wang, Xinyi Zhou, Lin Qiu, Joseph Chee Chang et al.CHI 2025 · 8 citations
- AMQuestioner: Training Critical Thinking with Question-Driven Interactive Argument Maps in Online DiscussionQiyu Pan, Jianqiao Zeng, Jie Wang, Junyu Liu et al.CSCW 2025 · 1 citation
- From Struggle to Success: Context-Aware Guidance for Screen Reader Users in Computer UseNan Chen, Jing Lu, Zilong Wang, Luna K. Qiu et al.CHI 2026 · 1 citation
Builds on31
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-trainingLinjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan et al.EMNLP 2020 · 387 citations
Related papers
- TutoAI: a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasksYuexi Chen, Vlad I. Morariu, Anh Truong, Zhicheng LiuCHI 2024 · 16 citations
- Supercharging Trial-and-Error for Learning Complex Software ApplicationsDamien Masson, Jo Vermeulen, George W. Fitzmaurice, Justin MatejkaCHI 2022 · 17 citations
- Answering Questions about Charts and Generating Visual ExplanationsDae Hyun Kim, Enamul Hoque, Maneesh AgrawalaCHI 2020 · 121 citations
- Automatic Instructional Video Creation from a Markdown-Formatted TutorialPeggy Chi, Nathan Frey, Katrina Panovich, Irfan EssaUIST 2021 · 27 citations
- YTCommentQA: Video Question Answerability in Instructional VideosSaelyne Yang, Sunghyun Park, Yunseok Jang, Moontae LeeAAAI 2024 · 6 citations
