LLMER: Crafting Interactive Extended Reality Worlds with JSON Data Generated by Large Language Models
Jiangong Chen, Xiaoyi Wu, Tian Lan, Bin Li
Abstract
The integration of Large Language Models (LLMs) like GPT-4 with Extended Reality (XR) technologies offers the potential to build truly immersive XR environments that interact with human users through natural language, e.g., generating and animating 3D scenes from audio inputs. However, the complexity of XR environments makes it difficult to accurately extract relevant contextual data and scene/object parameters from an overwhelming volume of XR artifacts. It leads to not only increased costs with pay-per-use models, but also elevated levels of generation errors. Moreover, existing approaches focusing on coding script generation are often prone to generation errors, resulting in flawed or invalid scripts, application crashes, and ultimately a degraded user experience. To overcome these challenges, we introduce LLMER, a novel framework that creates interactive XR worlds using JSON data generated by LLMs. Unlike prior approaches focusing on coding script generation, LLMER translates natural language inputs into JSON data, significantly reducing the likelihood of application crashes and processing latency. It employs a multi-stage strategy to supply only the essential contextual information adapted to the user's request and features multiple modules designed for various XR tasks. Our preliminary user study reveals the effectiveness of the proposed system, with over 80% reduction in consumed tokens and around 60% reduction in task completion time compared to state-of-the-art approaches. The analysis of users' feedback also illuminates a series of directions for further optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b89d458f-1aba-4789-ac07-2c50ae3989c3Cited by top-tier papers4
- Learning to Generate Structured Output with Schema Reinforcement LearningYaxi Lu, Haolun Li, Xin Cong, Zhong Zhang et al.ACL 2025 · 21 citations
- ImaginateAR: AI-Assisted In-Situ Authoring in Augmented RealityJaewook Lee, Filippo Aleotti, Diego Mazala, Guillermo Garcia-Hernando et al.UIST 2025 · 15 citations
- Roomify: Spatially-Grounded Style Transformation for Immersive Virtual EnvironmentsXueyang Wang, Qinxuan Cen, Weitao Bi, Yunxiang Ma et al.CHI 2026 · 1 citation
- PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language ModelsJiangong Chen, Mingyu Zhu, Bin LiIEEE VR 2026
Builds on9
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao et al.ICCV 2023 · 685 citations
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng et al.NeurIPS 2023 · 662 citations
Related papers
- LLMR: Real-time Prompting of Interactive Worlds using Large Language ModelsFernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski-Fahey et al.CHI 2024 · 124 citations
- GenAssist: Interactive Prompt-Driven XR Program GenerationSruti Srinidhi, Akul Singh, Edward Lu, Anthony RoweIEEE VR 2026
- LLM Integration in Extended Reality: A Comprehensive Review of Current Trends, Challenges, and Future PerspectivesYiliu Tang, Jason Situ, Andrea Yaoyun Cui, Mengke Wu et al.CHI 2025 · 37 citations
- Text2VRScene: Exploring the Framework of Automated Text-driven Generation System for VR ExperienceZhizhuo Yin, Yuyang Wang, Theodoros Papatheodorou, Pan HuiIEEE VR 2024 · 25 citations
- XRFix: Exploring Performance Bug Repair of Extended Reality Applications with Large Language ModelsJingwen Wu, Hanyang Guo, Hong-Ning Dai, Xiapu LuoICSE 2026
