Global-Aware Monocular Semantic Scene Completion with State Space Models
Shijie Li, Zhongyao Cheng, Rong Li, Shuai Li, Juergen Gall, Xun Xu, Xulei Yang
Abstract
Monocular Semantic Scene Completion (MonoSSC) reconstructs and interprets 3D environments from a single image, enabling diverse real-world applications. However, existing methods are often constrained by the local receptive field of Convolutional Neural Networks (CNNs), making it challenging to handle the non-uniform distribution of projected points (Fig. 1) and effectively reconstruct missing information caused by the 3D-to-2D projection. In this work, we introduce GA-MonoSSC, a hybrid architecture for MonoSSC that effectively captures global context in both the 2D image domain and 3D space. Specifically, we propose a Dual-Head Multi-Modality Encoder, which leverages a Transformer architecture to capture spatial relationships across all features in the 2D image domain, enabling more comprehensive 2D feature extraction. Additionally, we introduce the Frustum Mamba Decoder, built on the State Space Model (SSM), to efficiently capture long-range dependencies in 3D space. Furthermore, we propose a frustum reordering strategy within the Frustum Mamba Decoder to mitigate feature discontinuities in the reordered voxel sequence, ensuring better alignment with the scan mechanism of the State Space Model (SSM) for improved 3D representation learning. We conduct extensive experiments on the widely used Occ-ScanNet and NYUv2 datasets, demonstrating that our proposed method achieves state-of-the-art performance, validating its effectiveness. The code will be released upon acceptance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on18
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- PointMamba: A Simple State Space Model for Point Cloud AnalysisDingkang Liang, Xin Zhou, Wei Xu, Xingkui Zhu et al.NeurIPS 2024 · 380 citations
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 354 citations
- S4ND: Modeling Images and Videos as Multidimensional Signals with State SpacesEric Nguyen, Karan Goel, Albert Gu, Gordon W. Downs et al.NeurIPS 2022 · 267 citations
Related papers
- MonoScene: Monocular 3D Semantic Scene CompletionAnh-Quan Cao, Raoul de CharetteCVPR 2022 · 251 citations
- NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates SpaceJiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai et al.ICCV 2023 · 150 citations
- Skip Mamba Diffusion for Monocular 3D Semantic Scene CompletionLi Liang, Naveed Akhtar, Jordan Vice, Xiangrui Kong et al.AAAI 2025 · 10 citations
- DAPointMamba: Domain Adaptive Point Mamba for Point Cloud CompletionYinghui Li, Qianyu Zhou, Di Shao, Hao Yang et al.AAAI 2026 · 1 citation
- Pamba: Enhancing Global Interaction in Point Clouds via State Space ModelZhuoyuan Li, Yubo Ai, Jiahao Lu, Chuxin Wang et al.AAAI 2025 · 12 citations
