LLMER: Crafting Interactive Extended Reality Worlds with JSON Data Generated by Large Language Models
Jiangong Chen, Xiaoyi Wu, Tian Lan, Bin Li
摘要
The integration of Large Language Models (LLMs) like GPT-4 with Extended Reality (XR) technologies offers the potential to build truly immersive XR environments that interact with human users through natural language, e.g., generating and animating 3D scenes from audio inputs. However, the complexity of XR environments makes it difficult to accurately extract relevant contextual data and scene/object parameters from an overwhelming volume of XR artifacts. It leads to not only increased costs with pay-per-use models, but also elevated levels of generation errors. Moreover, existing approaches focusing on coding script generation are often prone to generation errors, resulting in flawed or invalid scripts, application crashes, and ultimately a degraded user experience. To overcome these challenges, we introduce LLMER, a novel framework that creates interactive XR worlds using JSON data generated by LLMs. Unlike prior approaches focusing on coding script generation, LLMER translates natural language inputs into JSON data, significantly reducing the likelihood of application crashes and processing latency. It employs a multi-stage strategy to supply only the essential contextual information adapted to the user's request and features multiple modules designed for various XR tasks. Our preliminary user study reveals the effectiveness of the proposed system, with over 80% reduction in consumed tokens and around 60% reduction in task completion time compared to state-of-the-art approaches. The analysis of users' feedback also illuminates a series of directions for further optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning to Generate Structured Output with Schema Reinforcement LearningYaxi Lu, Haolun Li, Xin Cong, Zhong Zhang 等ACL 2025 · 被引用 21 次
- ImaginateAR: AI-Assisted In-Situ Authoring in Augmented RealityJaewook Lee, Filippo Aleotti, Diego Mazala, Guillermo Garcia-Hernando 等UIST 2025 · 被引用 15 次
- Roomify: Spatially-Grounded Style Transformation for Immersive Virtual EnvironmentsXueyang Wang, Qinxuan Cen, Weitao Bi, Yunxiang Ma 等CHI 2026 · 被引用 1 次
- PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language ModelsJiangong Chen, Mingyu Zhu, Bin LiIEEE VR 2026
它引用的顶会 Paper9
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch 等ICML 2023 · 被引用 2,601 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language ModelsChan Hee Song, Brian M. Sadler, Jiaman Wu, Wei-Lun Chao 等ICCV 2023 · 被引用 685 次
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng 等NeurIPS 2023 · 被引用 662 次
相关 Paper
- LLMR: Real-time Prompting of Interactive Worlds using Large Language ModelsFernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski-Fahey 等CHI 2024 · 被引用 124 次
- GenAssist: Interactive Prompt-Driven XR Program GenerationSruti Srinidhi, Akul Singh, Edward Lu, Anthony RoweIEEE VR 2026
- LLM Integration in Extended Reality: A Comprehensive Review of Current Trends, Challenges, and Future PerspectivesYiliu Tang, Jason Situ, Andrea Yaoyun Cui, Mengke Wu 等CHI 2025 · 被引用 37 次
- Text2VRScene: Exploring the Framework of Automated Text-driven Generation System for VR ExperienceZhizhuo Yin, Yuyang Wang, Theodoros Papatheodorou, Pan HuiIEEE VR 2024 · 被引用 25 次
- XRFix: Exploring Performance Bug Repair of Extended Reality Applications with Large Language ModelsJingwen Wu, Hanyang Guo, Hong-Ning Dai, Xiapu LuoICSE 2026
