Interaction-Aware Shared Scene Synthesis for VR Telepresence
Zhangyao Tan, Qixiang Ma, Runze Fan, Sio Kei Im, Lili Wang
Abstract
Virtual reality telepresence requires immersive shared virtual environments for real-time remote collaboration across different physical scenes. It supports a wide range of applications in teleconferencing, education, and interactive simulations. However, challenges persist in identifying optimal shared virtual spaces that accommodate diverse user interaction requirements while adhering to local physical constraints during scene synthesis. In this paper, we propose an interaction-aware shared virtual scene synthesis method, which uses the large language model (LLM) to produce collaborative virtual scenes based on interaction demands from remote users in different local spaces. First, we introduce the concept of Interaction Aware Template (IAT) and its construction method using an LLM planner. Then, we propose an IAT-based affordance field alignment method for merging the local spaces of the remote users, maximally ensuring that the aligned space could support the user's desired interaction. Finally, we propose an LLM-based shared scene synthesis method according to the merged affordance field. Experiment results show that, compared to existing text-based scene synthesis and mutual space matching methods, our method achieves better Affordance Consistency, 3D Intersection over Union, and Layout Suitability on both scanned and synthesized datasets. The results of the user study demonstrate that the user's subjective perception of interaction fitness and sense of safety were significantly improved.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e9f058a-f8f8-45b1-9366-55352dea077bBuilds on34
- LayoutGPT: Compositional Visual Planning and Generation with Large Language ModelsWeixi Feng, Wanrong Zhu, Tsu-Jui Fu, Varun Jampani et al.NeurIPS 2023 · 462 citations
- 3D-FRONT: 3D Furnished Rooms with layOuts and semaNTicsHuan Fu, Bowen Cai, Lin Gao, Lingxiao Zhang et al.ICCV 2021 · 419 citations
- ATISS: Autoregressive Transformers for Indoor Scene SynthesisDespoina Paschalidou, Amlan Kar, Maria Shugrina, Karsten Kreis et al.NeurIPS 2021 · 293 citations
- Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image ModelsLukas Höllein, Ang Cao, Andrew Owens, Justin Johnson et al.ICCV 2023 · 292 citations
- RIO: 3D Object Instance Re-Localization in Changing Indoor EnvironmentsJohanna Wald, Armen Avetisyan, Nassir Navab, Federico Tombari et al.ICCV 2019 · 233 citations
Related papers
- Exploring Large Language Model-Driven Agents for Environment-Aware Spatial Interactions and Conversations in Virtual Reality Role-Play ScenariosZiming Li, Huadong Zhang, Chao Peng, Roshan L. PeirisIEEE VR 2025 · 18 citations
- HOG-Layout: Hierarchical 3D Scene Generation, Optimization and Editing via Vision-Language ModelsHaiyan Jiang, Deyu Zhang, Dongdong Weng, Weitao Song et al.CVPR 2026 · 1 citation
- In Situ 3D Scene Synthesis for Ubiquitous Embodied InterfacesHaiyan Jiang, Leiyu Song, Dongdong Weng, Zhe Sun et al.ACM MM 2024 · 3 citations
- MRUnion: Asymmetric Task-Aware 3D Mutual Scene Generation of Dissimilar Spaces for Mixed Reality TelepresenceMichael Pabst, Linda Rudolph, Nikolas Brasch, Verena Biener et al.IEEE VR 2025 · 8 citations
- Exploring Mediation by an Embodied Virtual Agent in Immersive Triadic Collaborative Decision-MakingBinyang Han, Ze Dong, Jingjing Zhang, Ruoyu Wen et al.IEEE VR 2026
