CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation
Li Liang, Bo Miao, Xinyu Wang, Naveed Akhtar, Jordan Vice, Ajmal Mian
摘要
Outdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in this direction are constrained by the absence of publicly available, well-annotated datasets. We introduce SketchSem3D, the first large-scale benchmark for generating 3D outdoor semantic scenes from abstract freehand sketches and pseudo-labeled annotations of satellite images. SketchSem3D includes two subsets, Sketch-based SemanticKITTI and Sketch-based KITTI-360 (containing LiDAR voxels along with their corresponding sketches and annotated satellite images), to enable standardized, rigorous, and diverse evaluations. We also propose Cylinder Mamba Diffusion (CymbaDiff) that significantly enhances spatial coherence in outdoor 3D scene generation. CymbaDiff imposes structured spatial ordering, explicitly captures cylindrical continuity and vertical hierarchy, and preserves both physical neighborhood relationships and global context within the generated scenes. Extensive experiments on SketchSem3D demonstrate that Cym-baDiff achieves superior semantic consistency, spatial realism, and cross-dataset generalization. The code and dataset will be available at https://github.com/Lillian- research-hub/CymbaDiff. scene layouts, diverse object geometries, and the need to preserve spatio-semantic coherence across large-scale scenes.
This work takes a significant step towards extending sketch-based generation to outdoor environments. To that end, we build upon the growth of State Space Models (SSMs) [24], which have gained increased attention across image segmentation [25,26] and point cloud processing [27,28] for their ability to capture long-range dependencies while remaining efficient through selective computation.However, to enhance global contextual understanding, SSMs typically aggregate information from multiple scan directions, leading to substantial memory overhead. Moreover, the scanning order imposed by the Cartesian coordinate system can distort local neighborhood relationships, especially in scenes with limited spatial coherence.
To address the above-noted challenges for sketch-based 3D outdoor scene generation, we first present 'SketchSem3D', a large-scale dataset tailored for the task. SketchSem3D enables the synthesis of semantically rich outdoor 3D environments from freehand sketches and pseudo-labeled satellite image annotations. The annotation pipeline properly integrates CLIP-based textual guidance [29] with image embeddings from the Segment Anything Model (SAM) [30], enabling robust and automated semantic labeling. SketchSem3D comprises two subsets, Sketch-based SemanticKITTI and Sketchbased KITTI-360, designed to support standardized benchmarking and fair comparison. Building upon this dataset, we define the novel 'sketch-based 3D outdoor scene generation' research task. We also propose Cylinder Mamba Diffusion, the first approach to handle this task. As adjacent Cartesian-based voxel sequences may misrepresent spatial proximity in outdoor scenes, CymbaDiff is particularly tailored to handle voxel discrepancies. Our underlying model is a denoising network, combining an SSM architecture with generative diffusion in the latent space. We design cylinder mamba blocks to enhance spatial coherence during the generative process, imposing a structured spatial ordering to explicitly encode cylindrical continuity and vertical hierarchy, preserving spatial neighborhood relationships within scenes.
Our key contributions are summarized below:
• We introduce the novel task of 'sketch-based 3D outdoor scene generation', which enables intuitive and flexible user interaction through freehand sketches and pseudo-labeled satellite image annotations. By reducing the need for manual semantic annotation, this task offers an efficient solution to generate training data for applications such as urban-scale simulation and autonomous driving.
• We present SketchSem3D, the first public large-scale sketch-based benchmark for 3D outdoor semantic scene generation. It includes two subsets, Sketch-based SemanticKITTI and Sketchbased KITTI-360, and enables standardized benchmarking for the development and evaluation of generative models in complex outdoor settings.
• We propose CymbaDiff, a generative model that incorporates the proposed cylinder mamba blocks to enhance spatial coherence during the generation process. We also conduct extensive experiments on the Sketch-based SemanticKITTI and Sketch-based KITTI-360 benchmarks, demonstrating state-of-the-art performance in 3D semantic scene generation and completion.
2 Related Work
Recent studies have demonstrated the strong capability of State-Space Models (SSMs) in capturing long-range dependencies across sequential data [31,32]. These models have been successfully applied in a variety of domains, including medical image segmentation [25,33], image restoration [34,35], natural language processing
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper56
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 被引用 3,482 次
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu 等NeurIPS 2024 · 被引用 3,199 次
相关 Paper
- Skip Mamba Diffusion for Monocular 3D Semantic Scene CompletionLi Liang, Naveed Akhtar, Jordan Vice, Xiangrui Kong 等AAAI 2025 · 被引用 10 次
- X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible ControllabilityYu Yang, Alan Liang, Jianbiao Mei, Yukai Ma 等NeurIPS 2025 · 被引用 22 次
- Controllable 3D Outdoor Scene Generation via Scene GraphsYuheng Liu, Xinke Li, Yuning Zhang, Lu Qi 等ICCV 2025 · 被引用 13 次
- Spiral: Semantic-Aware Progressive LiDAR Scene Generation and UnderstandingDekai Zhu, Yixuan Hu, Youquan Liu, Dongyue Lu 等NeurIPS 2025
- SemCity: Semantic Scene Generation with Triplane DiffusionJumin Lee, Sebin Lee, Changho Jo, Woobin Im 等CVPR 2024
