PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI
Yandan Yang, Baoxiong Jia, Peiyuan Zhi, Siyuan Huang
Abstract
With recent developments in Embodied Artificial Intel-ligence (EAI) research, there has been a growing demand for high-quality, large-scale interactive scene generation. While prior methods in scene synthesis have prioritized the naturalness and realism of the generated scenes, the physical plausibility and interactivity of scenes have been largely left unexplored. To address this disparity, we introduce PhyScene, a novel method dedicated to gener-ating interactive 3D scenes characterized by realistic lay-outs, articulated objects, and rich physical interactivity tai-lored for embodied agents. Based on a conditional diffusion model for capturing scene layouts, we devise novel physics-and interactivity-based guidance mechanisms that integrate constraints from object collision, room layout, and object reachability. Through extensive experiments, we demon-strate that PhyScene effectively leverages these guidance functions for physically interactable scene synthesis, out-performing existing state-of-the-art scene synthesis methods by a large margin. Our findings suggest that the scenes generated by PhyScene hold considerable potential for facilitating diverse skill acquisition among agents within in-teractive environments, thereby catalyzing further advance-ments in embodied AI research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1dad8425-a582-4018-b5b0-22c359704981Cited by top-tier papers44
- Sparc3D: Sparse Representation and Construction for High-Resolution 3D Shapes ModelingZhihao Li, Yufei Wang, Heliang Zheng, Yihao Luo et al.NeurIPS 2025 · 92 citations
- Scenethesis: A Language and Vision Agentic Framework for 3D Scene GenerationLu Ling, Chen-Hsuan Lin, Tsung-Yi Lin, Yifan Ding et al.ICLR 2026 · 74 citations
- SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective AgentYandan Yang, Baoxiong Jia, Shujie Zhang, Siyuan HuangNeurIPS 2025 · 65 citations
- PhyRecon: Physically Plausible Neural Scene ReconstructionJunfeng Ni, Yixin Chen, Bohan Jing, Nan Jiang et al.NeurIPS 2024 · 54 citations
- SAGE: Scalable Agentic 3D Scene Generation for Embodied AIHongchi Xia, Xuan Li, Zhaoshuo Li, Qianli Ma et al.CVPR 2026 · 50 citations
Builds on31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
- Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationJay Zhangjie Wu, Yixiao Ge, Xintao Wang, Stan Weixian Lei et al.ICCV 2023 · 1,113 citations
Related papers
- PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene ArrangementYian Wang, Han Yang, Minghao Guo, Xiaowen Qiu et al.ICLR 2026 · 10 citations
- Architect: Generating Vivid and Interactive 3D Scenes with Hierarchical 2D InpaintingYian Wang, Xiaowen Qiu, Jiageng Liu, Zhehuan Chen et al.NeurIPS 2024 · 48 citations
- PhyScene3D: Physically Consistent 3D Interactive Tabletop Scene GenerationWeixing Chen, Zhuoqian Feng, Yang Liu, Yexin Zhang et al.ICML 2026
- In Situ 3D Scene Synthesis for Ubiquitous Embodied InterfacesHaiyan Jiang, Leiyu Song, Dongdong Weng, Zhe Sun et al.ACM MM 2024 · 3 citations
- DynScene: Scalable Generation of Dynamic Robotic Manipulation Scenes for Embodied AISangmin Lee, Sungyong Park, Heewon KimCVPR 2025
