Guided Reality: Generating Visually-Enriched AR Task Guidance with LLMs and Vision Models
Ada Yi Zhao, Aditya Gunturu, Ellen Yi-Luen Do, Ryo Suzuki
Abstract
Large language models (LLMs) have enabled the automatic generation of step-by-step augmented reality (AR) instructions for a wide range of physical tasks. However, existing LLM-based AR guidance often lacks rich visual augmentations to effectively embed instructions into spatial context for a better user understanding. We present Guided Reality, a fully automated AR system that generates embedded and dynamic visual guidance based on step-by-step instructions. Our system integrates LLMs and vision models to: 1) generate multi-step instructions from user queries, 2) identify appropriate types of visual guidance, 3) extract spatial information about key interaction points in the real world, and 4) embed visual guidance in physical space to support task execution. Drawing from a corpus of user manuals, we define five categories of visual guidance and propose an identification strategy based on the current step. We evaluate the system through a user study (N=16), completing real-world tasks and exploring the system in the wild. Additionally, four instructors shared insights on how Guided Reality could be integrated into their training workflows.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9cda4e55-d07b-4a44-b3c1-452a78936c0fCited by top-tier papers2
- AgentHands: Generating Interactive Hand Gestures for Spatially Grounded Agent Conversations in XRZiyi Liu, David Li, Zhongyi Zhou, David Kim et al.CHI 2026 · 1 citation
- From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented RealityYoonsang Kim, Divyansh Pradhan, Devshree Jadeja, Arie E. KaufmanIEEE VR 2026
Builds on16
- A User Study on Mixed Reality Remote Collaboration with Eye Gaze and Hand Gesture SharingHuidong Bai, Prasanth Sasikumar, Jing Yang, Mark BillinghurstCHI 2020 · 210 citations
- AdapTutAR: An Adaptive Tutoring System for Machine Tasks in Augmented RealityGaoping Huang, Xun Qian, Tianyi Wang, Fagun Patel et al.CHI 2021 · 93 citations
- GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented RealityJaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu et al.CHI 2024 · 86 citations
- GesturAR: An Authoring System for Creating Freehand Interactive Augmented Reality ApplicationsTianyi Wang, Xun Qian, Fengming He, Xiyun Hu et al.UIST 2021 · 75 citations
- An Exploratory Study of Augmented Reality Presence for Tutoring Machine TasksYuanzhi Cao, Xun Qian, Tianyi Wang, Rachel Lee et al.CHI 2020 · 73 citations
Related papers
- AI-Powered Conversational Assistance in Augmented Reality for Multi-Step TasksJuliana H. Madritsch, Tomislav Duricic, Neven A. M. ElSayed, Simone Kopeinik et al.IEEE VR 2026 · 1 citation
- From Structure to Semantics: Hypergraph-Based AR Assembly Guidance with LLM-Mediated NarrationXinda Liu, Bowei Zhang, Jiaju Xu, Jian Wu et al.IEEE VR 2026
- Satori 悟り: Towards Proactive AR Assistant with Belief-Desire-Intention User ModelingChenyi Li, Guande Wu, Gromit Yeuk-Yin Chan, Dishita G. Turakhia et al.CHI 2025 · 49 citations
- Role-Aware Virtual Agents for Navigational Interaction guided by a Multimodal Large Language ModelMinyoung Kim, Changyang Li, Cuong Nguyen, Lap-Fai YuSIGGRAPH 2026
- AI-generated AR Reassembly Guidance from Disassembly Videos to Scaffold Everyday RepairWenjing Deng, Zhihao Yao, Xinhui Kang, Qirui Sun et al.CHI 2026 · 1 citation
