Video-Annotated Augmented Reality Assembly Tutorials
Masahiro Yamaguchi, Shohei Mori, Peter Mohr, Markus Tatzgern, Ana Stanescu, Hideo Saito, Denis Kalkofen
Abstract
We present a system for generating and visualizing interactive 3D Augmented Reality tutorials based on 2D video input, which allows viewpoint control at runtime. Inspired by assembly planning, we analyze the input video using a 3D CAD model of the object to determine an assembly graph that encodes blocking relationships between parts. Using an assembly graph enables us to detect assembly steps that are otherwise difficult to extract from the video, and generally improves object detection and tracking by providing prior knowledge about movable parts. To avoid information loss, we combine the 3D animation with relevant parts of the 2D video so that we can show detailed manipulations and tool usage that cannot be easily extracted from the video. To further support user orientation, we visually align the 3D animation with the real-world object by using texture information from the input video. We developed a presentation system that uses commonly available hardware to make our results accessible for home use and demonstrate the effectiveness of our approach by comparing it to traditional video tutorials.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Design Patterns for Situated Visualization in Augmented RealityBenjamin Lee, Michael Sedlmair, Dieter SchmalstiegIEEE VIS 2023 · 72 citations
- InstruMentAR: Auto-Generation of Augmented Reality Tutorials for Operating Digital Instruments Through Recording Embodied DemonstrationZiyi Liu, Zhengzhe Zhu, Enze Jiang, Feichi Huang et al.CHI 2023 · 29 citations
- PrISM-Tracker: A Framework for Multimodal Procedure Tracking Using Wearable Sensors and State Transition Information with User-Driven Handling of Errors and UncertaintyRiku Arakawa, Hiromu Yakura, Vimal Mollyn, Suzanne Nie et al.UbiComp 2023 · 20 citations
- PrISM-Observer: Intervention Agent to Help Users Perform Everyday Procedures Sensed using a SmartwatchRiku Arakawa, Hiromu Yakura, Mayank GoelUIST 2024 · 20 citations
- VoLearn: A Cross-Modal Operable Motion-Learning System Combined with Virtual Avatar and Auditory FeedbackChengshuo Xia, Xinrui Fang, Riku Arakawa, Yuta SugiuraUbiComp 2022 · 19 citations
Related papers
- Task Breakpoint Generation using Origin-Centric Graph in Virtual Reality Recordings for Adaptive PlaybackSelin Choi, Dooyoung Kim, Taewook Ha, Seonji Kim et al.IEEE VR 2026
- ARify: Leveraging Narrated Instructional Videos to Create Augmented Reality Tutorials for Procedural TasksXiyun Hu, Chenfei Zhu, Shao-Kang Hsia, Dizhi Ma et al.CHI 2026 · 1 citation
- Which Side is the Top? A User Study to Compare Visual Assets for Component Orientation in Assembly with Augmented RealityEnricoandrea Laviola, Michele Gattullo, Sara Romano, Antonio Emmanuele UvaIEEE VR 2025 · 2 citations
- SimRecon: SimReady Compositional Scene Reconstruction from Real VideosChong Xia, Kai Zhu, Zizhuo Wang, Fangfu Liu et al.CVPR 2026 · 11 citations
- Neural Assembler: Learning to Generate Fine-Grained Robotic Assembly Instructions from Multi-View ImagesHongyu Yan, Yadong MuAAAI 2025 · 3 citations
