SceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene Generation
Aleksei Bokhovkin, Quan Meng, Shubham Tulsiani, Angela Dai
摘要
Figure 1. SceneFactor factors the complex task of text-guided 3D scene generation into forming a coarse semantic structure, followed by refined geometric synthesis. Rather than require a learned model to decide the location, type, size, and local geometry of scene elements directly, our generation of a coarse semantic box layout enables training a simpler task of layout-guided geometric synthesis. To achieve this factorized generation, we train semantic and geometric latent diffusion models. Crucially, the proxy semantic map generation enables user-friendly localized editing of generated scenes by editing in the semantic map with simple box operations (by clicking two box corners), without requiring re-synthesis of the full scene. Note that input text is colored by semantic categories for visualization purposes only.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- WorldGen: From Text to Traversable and Interactive 3D WorldsDilin Wang, Hyunyoung Jung, Tom Monnier, Kihyuk Sohn 等CVPR 2026 · 被引用 24 次
- NuiScene: Exploring Efficient Generation of Unbounded Outdoor ScenesHan-Hung Lee, Qinghong Han, Angel X. ChangICCV 2025 · 被引用 5 次
- Attacks on Approximate Caches in Text-to-Image Diffusion ModelsDesen Sun, Shuncheng Jie, Sihang LiuUSENIX Security 2026 · 被引用 1 次
- FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial PriorsChenxi Li, Weijie Wang, Qiang Li, Nicu Sebe 等ACM MM 2025
- PrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic Urban ScenesChristina Ourania Tze, Daniel Dauner, Yiyi Liao, Dzmitry Tsishkou 等CVPR 2026
它引用的顶会 Paper45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- SceneGenesis: 3D Scene Synthesis via Semantic Structural Priors and Mesh-Guided Video-Geometry FusionYueming Zhao, Hongyu Yang, Di HuangAAAI 2026
- SceneCraft: Layout-Guided 3D Scene GenerationXiuyu Yang, Yunze Man, Jun-Kun Chen, Yu-Xiong WangNeurIPS 2024 · 被引用 58 次
- Towards Text-guided 3D Scene CompositionQihang Zhang, Chaoyang Wang, Aliaksandr Siarohin, Peiye Zhuang 等CVPR 2024 · 被引用 16 次
- Laconic: A 3D Layout Adapter for Controllable Image CreationLéopold Maillard, Tom Durand, Adrien Ramanana Rahary, Maks OvsjanikovICCV 2025
- Language-driven Scene Synthesis using Multi-conditional Diffusion ModelVuong Dinh An, Minh Nhat Vu, Toan Nguyen, Baoru Huang 等NeurIPS 2023 · 被引用 14 次
