Towards Text-guided 3D Scene Composition
Qihang Zhang, Chaoyang Wang, Aliaksandr Siarohin, Peiye Zhuang, Yinghao Xu, Ceyuan Yang, Dahua Lin, Bolei Zhou, Sergey Tulyakov, Hsin-Ying Lee
摘要
We are witnessing significant breakthroughs in the tech-nology for generating 3D objects from text. Existing approaches either leverage large text-to-image models to optimize a 3D representation or train 3D generators on object-centric datasets. Generating entire scenes, however, remains very challenging as a scene contains multiple 3D objects, diverse and scattered. In this work, we introduce SceneWiz3D - a novel approach to synthesize high-fidelity 3D scenes from text. We marry the locality of objects with globality of scenes by introducing a hybrid 3D representation - explicit for objects and implicit for scenes. Remarkably, an object, being represented explicitly, can be either generated from text using conventional text-to-3D approaches, or provided by users. To configure the layout of the scene and automatically place objects, we apply the Particle Swarm Optimization technique during the optimization process. Furthermore, it is difficult for certain parts of the scene (e.g., corners, occlusion) to receive multi-view supervision, leading to inferior geometry. We incor-porate an RGBD panorama diffusion model to mitigate it, resulting in high-quality geometry. Extensive evaluation supports that our approach achieves superior quality over previous approaches, enabling the generation of detailed and view-consistent 3D scenes. Our project website is at https://zqh0253.github.io/SceneWiz3D/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Direct Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial ReasoningXingjian Ran, Yixuan Li, Linning Xu, Mulin Yu 等NeurIPS 2025 · 被引用 34 次
- VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric ControlSixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li 等CVPR 2026 · 被引用 27 次
- InterGSEdit: Interactive 3D Gaussian Splatting Editing with 3D Geometry-Consistent Attention PriorMinghao Wen, Shengjie Wu, Kangkan Wang, Dong LiangICCV 2025 · 被引用 5 次
- HiScene: Creating Hierarchical 3D Scenes with Isometric View GenerationWenqi Dong, Bangbang Yang, Zesong Yang, Yuan Li 等ACM MM 2025 · 被引用 3 次
- MatLat: Material Latent Space for PBR Texture GenerationKyeongmin Yeo, Yunhong Min, Jaihoon Kim, Minhyuk SungCVPR 2026 · 被引用 2 次
它引用的顶会 Paper35
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Sculpt3D: Multi-View Consistent Text-to-3D Generation with Sparse 3D PriorCheng Chen, Xiaofeng Yang, Fan Yang, Chengzeng Feng 等CVPR 2024
- Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene GenerationYuanbo Yang, Jiahao Shao, Xinyang Li, Yujun Shen 等CVPR 2025
- 3D-SceneDreamer: Text-Driven 3D-Consistent Scene GenerationSongchun Zhang, Yibo Zhang, Quan Zheng, Rui Ma 等CVPR 2024 · 被引用 3 次
- SceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene GenerationAleksei Bokhovkin, Quan Meng, Shubham Tulsiani, Angela DaiCVPR 2025
- SceneGenesis: 3D Scene Synthesis via Semantic Structural Priors and Mesh-Guided Video-Geometry FusionYueming Zhao, Hongyu Yang, Di HuangAAAI 2026
