ImmerseGen: Agent-Guided Immersive World Generation with Alpha-Textured Proxies
Jinyan Yuan, Bangbang Yang, Keke Wang, Panwang Pan, Lin Ma, Xuehai Zhang, Xiao Liu, Zhaopeng Cui, Yuewen Ma
Abstract
Automating immersive VR scene creation remains a primary research challenge. Existing methods typically rely on complex geometry with post-simplification, resulting in inefficient pi pelines or li mited re alism. In th is paper, we introduce Im merseGen, a novel agent-guided framework for compact and photorealistic world generation that decouples realism from exhaustive geometric modeling. ImmerseGen represents scenes as hierarchical compositions of lightweight geometric proxies with synthesized RGBA textures, facilitating real-time rendering on mobile VR headsets. We propose terrain-conditioned texturing for base world generation, combined with context-aware texturing for scenery, to produce diverse and visually coherent worlds. VLM-based agents employ semantic grid-based analysis for precise asset placement and enrich scenes with multimodal enhancements such as visual dynamics and ambient sound. Experiments and real-time VR applications demonstrate that ImmerseGen achieves superior photorealism, spatial coherence, and rendering efficiency compared to existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f0fef5a-68d7-4d45-8997-b21023e1131eCited by top-tier papers3
- Code2Worlds: Empowering Coding LLMs for 4D World GenerationYi Zhang, Yunshuang Wang, Zeyu Zhang, Hao TangICML 2026 · 6 citations
- BuildAnyPoint: 3D Building Structured Abstraction from Diverse Point CloudsTongyan Hua, Haoran Gong, Yuan Liu, Di Wang et al.CVPR 2026 · 1 citation
- Diff4Splat: Repurposing Video Diffusion Models for Dynamic Scene GenerationPanwang Pan, Chenguo Lin, Chenxin Li, Jingjing Zhao et al.CVPR 2026
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
Related papers
- Text2Scene: Text-driven Indoor Scene Stylization with Part-Aware DetailsInwoo Hwang, Hyeonwoo Kim, Young Min KimCVPR 2023
- Generating Audiovisual Synergy Fluid Animation for Highly Immersive VR ExperienceNa Jiang, Xiangcheng Zhai, Yuxuan Qiu, Xiaohui Tan et al.IEEE VR 2026
- WorldGen: From Text to Traversable and Interactive 3D WorldsDilin Wang, Hyunyoung Jung, Tom Monnier, Kihyuk Sohn et al.CVPR 2026 · 24 citations
- HOG-Layout: Hierarchical 3D Scene Generation, Optimization and Editing via Vision-Language ModelsHaiyan Jiang, Deyu Zhang, Dongdong Weng, Weitao Song et al.CVPR 2026 · 1 citation
- ProXeek: Seeking and Leveraging Real-World Objects and Environments as Haptic Proxies for Virtual Reality through Multimodal ReasoningHaichen Gao, Tianrui Hu, Zeshui Li, Kening ZhuSIGGRAPH 2026
