Monocular Semantic Scene Completion via Masked Recurrent Networks
Xuzhi Wang, Xinran Wu, Song Wang, Lingdong Kong, Ziping Zhao
Abstract
Monocular Semantic Scene Completion (MSSC) aims to predict the voxel-wise occupancy and semantic category from a single-view RGB image. Existing methods adopt a single-stage framework that aims to simultaneously achieve visible region segmentation and occluded region hallucination, while also being affected by inaccurate depth estimation. Such methods often achieve suboptimal performance, especially in complex scenes. We propose a novel two-stage framework that decomposes MSSC into coarse MSSC followed by the Masked Recurrent Network. Specifically, we propose the Masked Sparse Gated Recurrent Unit (MS-GRU) which concentrates on the occupied regions by the proposed mask updating mechanism, and a sparse GRU design is proposed to reduce the computation cost. Additionally, we propose the distance attention projection to reduce projection errors by assigning different attention scores according to the distance to the observed surface. Experimental results demonstrate that our proposed unified framework, MonoMRN, effectively supports both indoor and outdoor scenes and achieves state-of-the-art performance on the NYUv2 and SemanticKITTI datasets. Furthermore, we conduct robustness analysis under various disturbances, highlighting the role of the Masked Recurrent Network in enhancing the model's resilience to such challenges. The source code is publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cff11d92-6e8d-43cc-82d0-1bd114dcf1cbCited by top-tier papers3
- OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial PerspectiveMarkus Gross, Sai B. Matha, Aya Fahmy, Rui Song et al.CVPR 2026 · 7 citations
- AdaSFormer: Adaptive Serialized Transformers for Monocular Semantic Scene Completion from Indoor EnvironmentsXuzhi Wang, Xinran Wu, Song Wang, Lingdong Kong et al.CVPR 2026 · 3 citations
- Spiral: Semantic-Aware Progressive LiDAR Scene Generation and UnderstandingDekai Zhu, Yixuan Hu, Youquan Liu, Dongyue Lu et al.NeurIPS 2025
Builds on48
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel et al.ICCV 2019 · 2,345 citations
- Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene CompletionXu Yan, Jiantao Gao, Jie Li, Ruimao Zhang et al.AAAI 2021 · 365 citations
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 354 citations
- MonoScene: Monocular 3D Semantic Scene CompletionAnh-Quan Cao, Raoul de CharetteCVPR 2022 · 251 citations
- Rethinking Range View Representation for LiDAR SegmentationLingdong Kong, Youquan Liu, Runnan Chen, Yuexin Ma et al.ICCV 2023 · 193 citations
Related papers
- Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene CompletionYu Xue, Longjun Gao, Yuanqi Su, HaoAng Lu et al.CVPR 2026
- Attention-Based Multi-Modal Fusion Network for Semantic Scene CompletionSiqi Li, Changqing Zou, Yipeng Li, Xibin Zhao et al.AAAI 2020 · 68 citations
- VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene CompletionYiming Li, Zhiding Yu, Christopher B. Choy, Chaowei Xiao et al.CVPR 2023
- Memory-Augmented Re-Completion for 3D Semantic Scene CompletionYu-Wen Tseng, Sheng-Ping Yang, Jhih-Ciang Wu, I-Bin Liao et al.AAAI 2025 · 3 citations
- NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates SpaceJiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai et al.ICCV 2023 · 150 citations
