Large Scene Generation with Cube-Absorb Discrete Diffusion
Qianjiang Hu Wei Hu, Wei Hu
摘要
Generating realistic 3D outdoor scenes is essential for applications in autonomous driving, virtual reality, environmental science, and urban development. Traditional 3D generation approaches using single-layer diffusion methods can produce detailed scenes for individual objects but struggle with high-resolution, large-scale outdoor environments due to scalability limitations. Recent hierarchical diffusion models tackle this by progressively scaling up lowresolution scenes. However, they often sample fine details from pure noise rather than from the coarse scene, which limits the efficiency. We propose a novel cubeabsorb discrete diffusion (CADD) model, which employs low-resolution scenes as the base state in the diffusion process to generate fine details, eliminating the need to sample entirely from noise. Moreover, we introduce the Sparse Cube Diffusion Transformer (SCDT), a transformer-based model with a sparse cube attention operator, optimized for generating large-scale sparse voxel scenes. Our method demonstrates state-of-the-art performance on the CarlaSC and KITTI360 datasets, supported by qualitative visualizations and extensive ablation studies that highlight the impact of the CADD process and sparse cube attention operator on high-resolution 3D scene generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper37
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- XCube: Large-Scale 3D Generative Modeling using Sparse Voxel HierarchiesXuanchi Ren, Jiahui Huang, Xiaohui Zeng, Ken Museth 等CVPR 2024 · 被引用 32 次
- Sat2City: 3D City Generation from a Single Satellite Image with Cascaded Latent DiffusionTongyan Hua, Lutao Jiang, Ying-Cong Chen, Wufan ZhaoICCV 2025 · 被引用 5 次
- SCube: Instant Large-Scale Scene Reconstruction using VoxSplatsXuanchi Ren, Yifan Lu, Hanxue Liang, Jay Zhangjie Wu 等NeurIPS 2024 · 被引用 63 次
- Sat2Scene: 3D Urban Scene Generation from Satellite Images with DiffusionZuoyue Li, Zhenqiang Li, Zhaopeng Cui, Marc Pollefeys 等CVPR 2024
- PrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic Urban ScenesChristina Ourania Tze, Daniel Dauner, Yiyi Liao, Dzmitry Tsishkou 等CVPR 2026
