SceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender Code
Ziniu Hu, Ahmet Iscen, Aashi Jain, Thomas Kipf, Yisong Yue, David A. Ross, Cordelia Schmid, Alireza Fathi
摘要
This paper introduces SceneCraft, a Large Language Model (LLM) Agent converting text descriptions into Blender-executable Python scripts which render complex scenes with up to a hundred 3D assets. This process requires complex spatial planning and arrangement. We tackle these challenges through a combination of advanced abstraction, strategic planning, and library learning. SceneCraft first models a scene graph as a blueprint, detailing the spatial relationships among assets in the scene. SceneCraft then writes Python scripts based on this graph, translating relationships into numerical constraints for asset layout. Next, SceneCraft leverages the perceptual strengths of vision-language foundation models like GPT-V to analyze rendered images and iteratively refine the scene. On top of this process, SceneCraft features a library learning mechanism that compiles common script functions into a reusable library, facilitating continuous self-improvement without expensive LLM parameter tuning. Our evaluation demonstrates that SceneCraft surpasses existing LLM-based agents in rendering complex scenes, as shown by its adherence to constraints and favorable human assessments. We also showcase the broader application potential of SceneCraft by reconstructing detailed 3D scenes from the Sintel movie and guiding a video generative model with generated scenes as intermediary control signal.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper34
- Scenethesis: A Language and Vision Agentic Framework for 3D Scene GenerationLu Ling, Chen-Hsuan Lin, Tsung-Yi Lin, Yifan Ding 等ICLR 2026 · 被引用 74 次
- WorldGen: From Text to Traversable and Interactive 3D WorldsDilin Wang, Hyunyoung Jung, Tom Monnier, Kihyuk Sohn 等CVPR 2026 · 被引用 24 次
- SceneGenAgent: Precise Industrial Scene Generation with Coding AgentXiao Xia, Dan Zhang, Zibo Liao, Zhenyu Hou 等ACL 2025 · 被引用 12 次
- Virtual Community: An Open World for Humans, Robots, and SocietyQinhong Zhou, Hongxin Zhang, Xiangye Lin, Zheyuan Zhang 等ICLR 2026 · 被引用 12 次
- MoVer: Motion Verification for Motion Graphics AnimationsJiaju Ma, Maneesh AgrawalaSIGGRAPH 2025 · 被引用 8 次
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- AgentBench: Evaluating LLMs as AgentsXiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu 等ICLR 2024 · 被引用 748 次
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang 等ICLR 2024 · 被引用 582 次
- GPT-4V(ision) is a Generalist Web Agent, if GroundedBoyuan Zheng, Boyu Gou, Jihyung Kil, Huan Sun 等ICML 2024 · 被引用 496 次
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama 等ICML 2024 · 被引用 464 次
相关 Paper
- ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D ModelingShuyuan Zhang, Chenhan Jiang, Zuoou Li, Jiankang DengNeurIPS 2025 · 被引用 6 次
- SceneCraft: Layout-Guided 3D Scene GenerationXiuyu Yang, Yunze Man, Jun-Kun Chen, Yu-Xiong WangNeurIPS 2024 · 被引用 58 次
- LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object IntegrationYuyao Zhang, Jinghao Li, Yu-Wing TaiNeurIPS 2025 · 被引用 21 次
- SceneX: Procedural Controllable Large-Scale Scene GenerationMengqi Zhou, Yuxi Wang, Jun Hou, Shougao Zhang 等AAAI 2025 · 被引用 22 次
- SceneGenesis: 3D Scene Synthesis via Semantic Structural Priors and Mesh-Guided Video-Geometry FusionYueming Zhao, Hongyu Yang, Di HuangAAAI 2026
