CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation
Li Liang, Bo Miao, Xinyu Wang, Naveed Akhtar, Jordan Vice, Ajmal Mian
Abstract
Outdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in this direction are constrained by the absence of publicly available, well-annotated datasets. We introduce SketchSem3D, the first large-scale benchmark for generating 3D outdoor semantic scenes from abstract freehand sketches and pseudo-labeled annotations of satellite images. SketchSem3D includes two subsets, Sketch-based SemanticKITTI and Sketch-based KITTI-360 (containing LiDAR voxels along with their corresponding sketches and annotated satellite images), to enable standardized, rigorous, and diverse evaluations. We also propose Cylinder Mamba Diffusion (CymbaDiff) that significantly enhances spatial coherence in outdoor 3D scene generation. CymbaDiff imposes structured spatial ordering, explicitly captures cylindrical continuity and vertical hierarchy, and preserves both physical neighborhood relationships and global context within the generated scenes. Extensive experiments on SketchSem3D demonstrate that Cym-baDiff achieves superior semantic consistency, spatial realism, and cross-dataset generalization. The code and dataset will be available at https://github.com/Lillian- research-hub/CymbaDiff. scene layouts, diverse object geometries, and the need to preserve spatio-semantic coherence across large-scale scenes.
This work takes a significant step towards extending sketch-based generation to outdoor environments. To that end, we build upon the growth of State Space Models (SSMs) [24], which have gained increased attention across image segmentation [25,26] and point cloud processing [27,28] for their ability to capture long-range dependencies while remaining efficient through selective computation.However, to enhance global contextual understanding, SSMs typically aggregate information from multiple scan directions, leading to substantial memory overhead. Moreover, the scanning order imposed by the Cartesian coordinate system can distort local neighborhood relationships, especially in scenes with limited spatial coherence.
To address the above-noted challenges for sketch-based 3D outdoor scene generation, we first present 'SketchSem3D', a large-scale dataset tailored for the task. SketchSem3D enables the synthesis of semantically rich outdoor 3D environments from freehand sketches and pseudo-labeled satellite image annotations. The annotation pipeline properly integrates CLIP-based textual guidance [29] with image embeddings from the Segment Anything Model (SAM) [30], enabling robust and automated semantic labeling. SketchSem3D comprises two subsets, Sketch-based SemanticKITTI and Sketchbased KITTI-360, designed to support standardized benchmarking and fair comparison. Building upon this dataset, we define the novel 'sketch-based 3D outdoor scene generation' research task. We also propose Cylinder Mamba Diffusion, the first approach to handle this task. As adjacent Cartesian-based voxel sequences may misrepresent spatial proximity in outdoor scenes, CymbaDiff is particularly tailored to handle voxel discrepancies. Our underlying model is a denoising network, combining an SSM architecture with generative diffusion in the latent space. We design cylinder mamba blocks to enhance spatial coherence during the generative process, imposing a structured spatial ordering to explicitly encode cylindrical continuity and vertical hierarchy, preserving spatial neighborhood relationships within scenes.
Our key contributions are summarized below:
• We introduce the novel task of 'sketch-based 3D outdoor scene generation', which enables intuitive and flexible user interaction through freehand sketches and pseudo-labeled satellite image annotations. By reducing the need for manual semantic annotation, this task offers an efficient solution to generate training data for applications such as urban-scale simulation and autonomous driving.
• We present SketchSem3D, the first public large-scale sketch-based benchmark for 3D outdoor semantic scene generation. It includes two subsets, Sketch-based SemanticKITTI and Sketchbased KITTI-360, and enables standardized benchmarking for the development and evaluation of generative models in complex outdoor settings.
• We propose CymbaDiff, a generative model that incorporates the proposed cylinder mamba blocks to enhance spatial coherence during the generation process. We also conduct extensive experiments on the Sketch-based SemanticKITTI and Sketch-based KITTI-360 benchmarks, demonstrating state-of-the-art performance in 3D semantic scene generation and completion.
2 Related Work
Recent studies have demonstrated the strong capability of State-Space Models (SSMs) in capturing long-range dependencies across sequential data [31,32]. These models have been successfully applied in a variety of domains, including medical image segmentation [25,33], image restoration [34,35], natural language processing
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2449e92e-72ab-4a06-b29b-cb40dbab8fc9Cited by top-tier papers1
Ask how each one uses itBuilds on56
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- VMamba: Visual State Space ModelYue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu et al.NeurIPS 2024 · 3,199 citations
Related papers
- Skip Mamba Diffusion for Monocular 3D Semantic Scene CompletionLi Liang, Naveed Akhtar, Jordan Vice, Xiangrui Kong et al.AAAI 2025 · 10 citations
- X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible ControllabilityYu Yang, Alan Liang, Jianbiao Mei, Yukai Ma et al.NeurIPS 2025 · 22 citations
- Controllable 3D Outdoor Scene Generation via Scene GraphsYuheng Liu, Xinke Li, Yuning Zhang, Lu Qi et al.ICCV 2025 · 13 citations
- Spiral: Semantic-Aware Progressive LiDAR Scene Generation and UnderstandingDekai Zhu, Yixuan Hu, Youquan Liu, Dongyue Lu et al.NeurIPS 2025
- SemCity: Semantic Scene Generation with Triplane DiffusionJumin Lee, Sebin Lee, Changho Jo, Woobin Im et al.CVPR 2024
