PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement
Yian Wang, Han Yang, Minghao Guo, Xiaowen Qiu, Tsun-Hsuan Johnson Wang, Wojciech Matusik, Joshua B. Tenenbaum, Chuang Gan
Abstract
Automatically generating interactive 3D environments is crucial for scaling up robotic data collection in simulation. While prior work has primarily focused on 3D asset placement, it often overlooks the physical relationships between objects (e.g., contact, support, balance, and containment), which are essential for creating complex and realistic manipulation scenarios such as tabletop arrangements, shelf organization, or box packing. Compared to classical 3D layout generation, producing complex physical scenes introduces additional challenges: (a) higher object density and complexity (e.g., a small shelf may hold dozens of books), (b) richer supporting relationships and compact spatial layouts, and (c) the need to accurately model both spatial placement and physical properties. To address these challenges, we propose PhyScensis, an LLM agent-based framework powered by a physics engine, to produce physically plausible scene configurations with high complexity. Specifically, our framework consists of three main components: an LLM agent iteratively proposes assets with spatial and physical predicates; a solver, equipped with a physics engine, realizes these predicates into a 3D scene; and feedback from the solver informs the agent to refine and enrich the configuration. Moreover, our framework preserves strong controllability over fine-grained textual descriptions and numerical parameters (e.g., relative positions, scene stability), enabled through probabilistic programming for stability and a complementary heuristic that jointly regulates stability and spatial relations. Experimental results show that our method outperforms prior approaches in scene complexity, visual quality, and physical accuracy, offering a unified pipeline for generating complex physical scene layouts for robotic manipulation. More qualitative results are on physcensis.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cdb1aa96-cc1c-4139-b776-93b9bb120636Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
Related papers
- Scenethesis: A Language and Vision Agentic Framework for 3D Scene GenerationLu Ling, Chen-Hsuan Lin, Tsung-Yi Lin, Yifan Ding et al.ICLR 2026 · 74 citations
- PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AIYandan Yang, Baoxiong Jia, Peiyuan Zhi, Siyuan HuangCVPR 2024 · 27 citations
- PAT3D: Physics-Augmented Text-to-3D Scene GenerationGuying Lin, Kemeng Huang, Michael Liu, Ruihan Gao et al.ICLR 2026 · 14 citations
- Pair2Scene: Learning Local Object Relations for Procedural Scene GenerationXingjian Ran, Shujie Zhang, Weipeng Zhong, Luo Li et al.ICML 2026
- SceneSmith: Agentic Generation of Simulation-Ready Indoor ScenesNicholas Pfaff, Thomas Cohn, Sergey Zakharov, Rick Cory et al.ICML 2026 · 21 citations
