Lune

NeurIPS2025顶会

CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation

Li Liang, Bo Miao, Xinyu Wang, Naveed Akhtar, Jordan Vice, Ajmal Mian

2025年份
4被引次数
1顶会引用

摘要

Outdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in this direction are constrained by the absence of publicly available, well-annotated datasets. We introduce SketchSem3D, the first large-scale benchmark for generating 3D outdoor semantic scenes from abstract freehand sketches and pseudo-labeled annotations of satellite images. SketchSem3D includes two subsets, Sketch-based SemanticKITTI and Sketch-based KITTI-360 (containing LiDAR voxels along with their corresponding sketches and annotated satellite images), to enable standardized, rigorous, and diverse evaluations. We also propose Cylinder Mamba Diffusion (CymbaDiff) that significantly enhances spatial coherence in outdoor 3D scene generation. CymbaDiff imposes structured spatial ordering, explicitly captures cylindrical continuity and vertical hierarchy, and preserves both physical neighborhood relationships and global context within the generated scenes. Extensive experiments on SketchSem3D demonstrate that Cym-baDiff achieves superior semantic consistency, spatial realism, and cross-dataset generalization. The code and dataset will be available at https://github.com/Lillian- research-hub/CymbaDiff. scene layouts, diverse object geometries, and the need to preserve spatio-semantic coherence across large-scale scenes.

This work takes a significant step towards extending sketch-based generation to outdoor environments. To that end, we build upon the growth of State Space Models (SSMs) [24], which have gained increased attention across image segmentation [25,26] and point cloud processing [27,28] for their ability to capture long-range dependencies while remaining efficient through selective computation.However, to enhance global contextual understanding, SSMs typically aggregate information from multiple scan directions, leading to substantial memory overhead. Moreover, the scanning order imposed by the Cartesian coordinate system can distort local neighborhood relationships, especially in scenes with limited spatial coherence.

To address the above-noted challenges for sketch-based 3D outdoor scene generation, we first present 'SketchSem3D', a large-scale dataset tailored for the task. SketchSem3D enables the synthesis of semantically rich outdoor 3D environments from freehand sketches and pseudo-labeled satellite image annotations. The annotation pipeline properly integrates CLIP-based textual guidance [29] with image embeddings from the Segment Anything Model (SAM) [30], enabling robust and automated semantic labeling. SketchSem3D comprises two subsets, Sketch-based SemanticKITTI and Sketchbased KITTI-360, designed to support standardized benchmarking and fair comparison. Building upon this dataset, we define the novel 'sketch-based 3D outdoor scene generation' research task. We also propose Cylinder Mamba Diffusion, the first approach to handle this task. As adjacent Cartesian-based voxel sequences may misrepresent spatial proximity in outdoor scenes, CymbaDiff is particularly tailored to handle voxel discrepancies. Our underlying model is a denoising network, combining an SSM architecture with generative diffusion in the latent space. We design cylinder mamba blocks to enhance spatial coherence during the generative process, imposing a structured spatial ordering to explicitly encode cylindrical continuity and vertical hierarchy, preserving spatial neighborhood relationships within scenes.

Our key contributions are summarized below:

• We introduce the novel task of 'sketch-based 3D outdoor scene generation', which enables intuitive and flexible user interaction through freehand sketches and pseudo-labeled satellite image annotations. By reducing the need for manual semantic annotation, this task offers an efficient solution to generate training data for applications such as urban-scale simulation and autonomous driving.

• We present SketchSem3D, the first public large-scale sketch-based benchmark for 3D outdoor semantic scene generation. It includes two subsets, Sketch-based SemanticKITTI and Sketchbased KITTI-360, and enables standardized benchmarking for the development and evaluation of generative models in complex outdoor settings.

• We propose CymbaDiff, a generative model that incorporates the proposed cylinder mamba blocks to enhance spatial coherence during the generation process. We also conduct extensive experiments on the Sketch-based SemanticKITTI and Sketch-based KITTI-360 benchmarks, demonstrating state-of-the-art performance in 3D semantic scene generation and completion.

2 Related Work

Recent studies have demonstrated the strong capability of State-Space Models (SSMs) in capturing long-range dependencies across sequential data [31,32]. These models have been successfully applied in a variety of domains, including medical image segmentation [25,33], image restoration [34,35], natural language processing

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper56

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖