Text2VRScene: Exploring the Framework of Automated Text-driven Generation System for VR Experience
Zhizhuo Yin, Yuyang Wang, Theodoros Papatheodorou, Pan Hui
Abstract
With the recent development of the Virtual Reality (VR) industry, the increasing number of VR users pushes the demand for the massive production of immersive and expressive VR scenes in related industries. However, creating expressive VR scenes involves the reasonable organization of various digital content to express a coherent and logical theme, which is time-consuming and labor-intensive. In recent years, Large Language Models (LLMs) such as ChatGPT 3.5 and generative models such as stable diffusion have emerged as powerful tools for comprehending natural language and generating digital contents such as text, code, images, and 3D objects. In this paper, we have explored how we can generate VR scenes from text by incorporating LLMs and various generative models into an automated system. To achieve this, we first identify the possible limitations of LLMs for an automated system and propose a systematic framework to mitigate them. Subsequently, we developed Text2VRScene, a VR scene generation system, based on our proposed framework with well-designed prompts. To validate the effectiveness of our proposed framework and the designed prompts, we carry out a series of test cases. The results show that the proposed framework contributes to improving the reliability of the system and the quality of the generated VR scenes. The results also illustrate the promising performance of the Text2VRScene in generating satisfying VR scenes with a clear theme regularized by our well-designed prompts. This paper ends with a discussion about the limitations of the current system and the potential of developing similar generation systems based on our framework.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers7
- Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System DesignYongquan 'Owen' Hu, Jingyu Tang, Xinya Gong, Zhongyi Zhou et al.CHI 2025 · 37 citations
- Semantics-Controlled Gaussian Splatting for Outdoor Scene Reconstruction and Rendering in Virtual RealityHannah Schieber, Jacob Young, Tobias Langlotz, Stefanie Zollmann et al.IEEE VR 2025 · 10 citations
- HOICraft: In-Situ VLM-based Authoring Tool for Part-Level Hand-Object Interaction Design in VRDohui Lee, Qi Sun, Sang Ho YoonCHI 2026 · 1 citation
- TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse RouteHongyi Luo, Qing Cheng, Daniel Matos, Hari Krishna Gadi et al.EMNLP 2025
- Interaction-Aware Shared Scene Synthesis for VR TelepresenceZhangyao Tan, Qixiang Ma, Runze Fan, Sio Kei Im et al.IEEE VR 2026
Related papers
- WorldGen: From Text to Traversable and Interactive 3D WorldsDilin Wang, Hyunyoung Jung, Tom Monnier, Kihyuk Sohn et al.CVPR 2026 · 24 citations
- Exploring Large Language Model-Driven Agents for Environment-Aware Spatial Interactions and Conversations in Virtual Reality Role-Play ScenariosZiming Li, Huadong Zhang, Chao Peng, Roshan L. PeirisIEEE VR 2025 · 18 citations
- GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian SplattingXiaoyu Zhou, Xingjian Ran, Yajiao Xiong, Jinlin He et al.ICML 2024 · 113 citations
- LLMER: Crafting Interactive Extended Reality Worlds with JSON Data Generated by Large Language ModelsJiangong Chen, Xiaoyi Wu, Tian Lan, Bin LiIEEE VR 2025 · 20 citations
- Analyzing Multimodal Interaction Strategies for LLM-Assisted Manipulation of 3D ScenesJunlong Chen, Jens Grubert, Per Ola KristenssonIEEE VR 2025 · 9 citations
