Soloist: Generating Mixed-Initiative Tutorials from Existing Guitar Instructional Videos Through Audio Processing
Bryan Wang, Mengyu Yang, Tovi Grossman
Abstract
Learning musical instruments using online instructional videos has become increasingly prevalent. However, pre-recorded videos lack the instantaneous feedback and personal tailoring that human tutors provide. In addition, existing video navigations are not optimized for instrument learning, making the learning experience encumbered. Guided by our formative interviews with guitar players and prior literature, we designed Soloist, a mixed-initiative learning framework that automatically generates customizable curriculums from off-the-shelf guitar video lessons. Soloist takes raw videos as input and leverages deep-learning based audio processing to extract musical information. This back-end processing is used to provide an interactive visualization to support effective video navigation and real-time feedback on the user's performance, creating a guided learning experience. We demonstrate the capabilities and specific use-cases of Soloist within the domain of learning electric guitar solos using instructional YouTube videos. A remote user study, conducted to gather feedback from guitar players, shows encouraging results as the users unanimously preferred learning with Soloist over unconverted instructional videos.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- A Literature Review of Video-Sharing Platform Research in HCIAva Bartolome, Shuo NiuCHI 2023 · 63 citations
- Synthesis-Assisted Video Prototyping From a DocumentPeggy Chi, Tao Dong, Christian Früh, Brian Colonna et al.UIST 2022 · 18 citations
- TutoAI: a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasksYuexi Chen, Vlad I. Morariu, Anh Truong, Zhicheng LiuCHI 2024 · 16 citations
- Video2Action: Reducing Human Interactions in Action Annotation of App Tutorial VideosSidong Feng, Chunyang Chen, Zhenchang XingUIST 2023 · 12 citations
- NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video UnderstandingRunning Zhao, Zhihan Jiang, Xinchen Zhang, Chirui Chang et al.UIST 2025 · 6 citations
Builds on2
- Temporal Segmentation of Creative Live StreamsC. Ailie Fraser, Joy O. Kim, Hijung Valentina Shin, Joel Brandt et al.CHI 2020 · 38 citations
- Exploring the Potential of an Intelligent Tutoring System for Sketching FundamentalsBlake Williford, Matthew Runyon, Wayne Li, Julie Linsey et al.CHI 2020 · 20 citations
Related papers
- The Sound of MotionsHang Zhao, Chuang Gan, Wei-Chiu Ma, Antonio TorralbaICCV 2019 · 271 citations
- ReTouche: Embodied Representations for Self-Guided Piano LearningPaul-Peter Arslan, Hayoun Noh, Mariana Aki Tamashiro, Louis Badr et al.CHI 2026 · 1 citation
- Music Gesture for Visual Sound SeparationChuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum et al.CVPR 2020
- Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric VideosLuigi Seminara, Giovanni Maria Farinella, Antonino FurnariNeurIPS 2024 · 36 citations
- Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen SoundsEfthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey et al.ICLR 2021 · 83 citations
