Extend3D: Town-Scale 3D Generation
Seungwoo Yoon, Jinmo Kim, Jaesik Park
摘要
In this paper, we propose Extend3D, a novel training-free pipeline for 3D scene generation from a single image, built upon an object-centric 3D generative model. To overcome the limitations of fixed-size latent spaces of object-centric models in representing wide scenes, we extend the latent space in and directions. Then, by dividing the extended latent into overlapping patches, we use the object-centric 3D generative model on each patch and couple them at each time step. Since object-centric models are sub-optimal for sub-scene generation, we use the input image and point cloud extracted from a depth estimator as priors to enable this process. Using the point cloud prior, we initialize the scene structure and refine the occluded region iteratively with under-noised SDEdit. Also, both priors are used to optimize the extended latent during the denoising process so that the denoising paths do not deviate from the sub-scene dynamics. We demonstrate that our method produces better results than previous methods, as evidenced by human preferences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- ForkNet: Multi-Branch Volumetric Semantic Completion From a Single Depth ImageYida Wang, David Joseph Tan, Nassir Navab, Federico TombariICCV 2019 · 被引用 67 次
- Unsupervised Causal Generative Understanding of ImagesTitas Anciukevicius, Patrick Fox-Roberts, Edward Rosten, Paul HendersonNeurIPS 2022 · 被引用 6 次
- Baking Gaussian Splatting Into Diffusion Denoiser for Fast and Scalable Single-Stage Image-to-3D Generation and ReconstructionYuanhao Cai, He Zhang, Kai Zhang, Yixun Liang 等ICCV 2025 · 被引用 7 次
- BlockGAN: Learning 3D Object-aware Scene Representations from Unlabelled ImagesThu Nguyen-Phuoc, Christian Richardt, Long Mai, Yong-Liang Yang 等NeurIPS 2020 · 被引用 256 次
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single ImageZe-Xin Yin, Liu Liu, Xinjie wang, Wei Sui 等CVPR 2026 · 被引用 10 次
