Skip Mamba Diffusion for Monocular 3D Semantic Scene Completion
Li Liang, Naveed Akhtar, Jordan Vice, Xiangrui Kong, Ajmal Saeed Mian
Abstract
3D semantic scene completion is critical for multiple downstream tasks in autonomous systems. It estimates missing geometric and semantic information in the acquired scene data. Due to the challenging real-world conditions, this task usually demands complex models that process multi-modal data to achieve acceptable performance. We propose a unique neural model, leveraging advances from the state space and diffusion generative modeling to achieve remarkable 3D semantic scene completion performance with monocular image input. Our technique processes the data in the conditioned latent space of a variational autoencoder where diffusion modeling is carried out with an innovative state space technique. A key component of our neural network is the proposed Skimba (Skip Mamba) denoiser, which is adept at efficiently processing long-sequence data. The Skimba diffusion model is integral to our 3D scene completion network, incorporating a triple Mamba structure, dimensional decomposition residuals and varying dilations along three directions. We also adopt a variant of this network for the subsequent semantic segmentation stage of our method. Extensive evaluation on the standard SemanticKITTI and SSCBench-KITTI360 datasets show that our approach not only outperforms other monocular techniques by a large margin, it also achieves competitive performance against stereo methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91b97260-cc1d-43ce-9a24-4608d95d8906Cited by top-tier papers5
- CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene GenerationLi Liang, Bo Miao, Xinyu Wang, Naveed Akhtar et al.NeurIPS 2025 · 4 citations
- VoxDet: Rethinking 3D Semantic Scene Completion as Dense Object DetectionWuyang Li, Zhu Yu, Alexandre AlahiNeurIPS 2025 · 3 citations
- Global-Aware Monocular Semantic Scene Completion with State Space ModelsShijie Li, Zhongyao Cheng, Rong Li, Shuai Li et al.ICCV 2025 · 1 citation
- RI-Mamba: Rotation-Invariant Mamba for Robust Text-to-Shape RetrievalKhanh Nguyen, Dasith de Silva Edirimuni, Ghulam Mubashar Hassan, Ajmal MianCVPR 2026
- Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene CompletionYu Xue, Longjun Gao, Yuanqi Su, HaoAng Lu et al.CVPR 2026
Builds on34
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
Related papers
- MeshMamba: State Space Models for Articulated 3D Mesh Generation and ReconstructionYusuke Yoshiyasu, Leyuan Sun, Ryusuke SagawaICCV 2025 · 2 citations
- MonoScene: Monocular 3D Semantic Scene CompletionAnh-Quan Cao, Raoul de CharetteCVPR 2022 · 251 citations
- ForkNet: Multi-Branch Volumetric Semantic Completion From a Single Depth ImageYida Wang, David Joseph Tan, Nassir Navab, Federico TombariICCV 2019 · 67 citations
- NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates SpaceJiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai et al.ICCV 2023 · 150 citations
- Disentangling Instance and Scene Contexts for 3D Semantic Scene CompletionEnyu Liu, En Yu, Sijia Chen, Wenbing TaoICCV 2025 · 2 citations
