Monocular Semantic Scene Completion via Masked Recurrent Networks
Xuzhi Wang, Xinran Wu, Song Wang, Lingdong Kong, Ziping Zhao
摘要
Monocular Semantic Scene Completion (MSSC) aims to predict the voxel-wise occupancy and semantic category from a single-view RGB image. Existing methods adopt a single-stage framework that aims to simultaneously achieve visible region segmentation and occluded region hallucination, while also being affected by inaccurate depth estimation. Such methods often achieve suboptimal performance, especially in complex scenes. We propose a novel two-stage framework that decomposes MSSC into coarse MSSC followed by the Masked Recurrent Network. Specifically, we propose the Masked Sparse Gated Recurrent Unit (MS-GRU) which concentrates on the occupied regions by the proposed mask updating mechanism, and a sparse GRU design is proposed to reduce the computation cost. Additionally, we propose the distance attention projection to reduce projection errors by assigning different attention scores according to the distance to the observed surface. Experimental results demonstrate that our proposed unified framework, MonoMRN, effectively supports both indoor and outdoor scenes and achieves state-of-the-art performance on the NYUv2 and SemanticKITTI datasets. Furthermore, we conduct robustness analysis under various disturbances, highlighting the role of the Masked Recurrent Network in enhancing the model's resilience to such challenges. The source code is publicly available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial PerspectiveMarkus Gross, Sai B. Matha, Aya Fahmy, Rui Song 等CVPR 2026 · 被引用 7 次
- AdaSFormer: Adaptive Serialized Transformers for Monocular Semantic Scene Completion from Indoor EnvironmentsXuzhi Wang, Xinran Wu, Song Wang, Lingdong Kong 等CVPR 2026 · 被引用 3 次
- Spiral: Semantic-Aware Progressive LiDAR Scene Generation and UnderstandingDekai Zhu, Yixuan Hu, Youquan Liu, Dongyue Lu 等NeurIPS 2025
它引用的顶会 Paper48
- SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR SequencesJens Behley, Martin Garbade, Andres Milioto, Jan Quenzel 等ICCV 2019 · 被引用 2,345 次
- Sparse Single Sweep LiDAR Point Cloud Segmentation via Learning Contextual Shape Priors from Scene CompletionXu Yan, Jiantao Gao, Jie Li, Ruimao Zhang 等AAAI 2021 · 被引用 365 次
- OccFormer: Dual-path Transformer for Vision-based 3D Semantic Occupancy PredictionYunpeng Zhang, Zheng Zhu, Dalong DuICCV 2023 · 被引用 354 次
- MonoScene: Monocular 3D Semantic Scene CompletionAnh-Quan Cao, Raoul de CharetteCVPR 2022 · 被引用 251 次
- Rethinking Range View Representation for LiDAR SegmentationLingdong Kong, Youquan Liu, Runnan Chen, Yuexin Ma 等ICCV 2023 · 被引用 193 次
相关 Paper
- Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene CompletionYu Xue, Longjun Gao, Yuanqi Su, HaoAng Lu 等CVPR 2026
- Attention-Based Multi-Modal Fusion Network for Semantic Scene CompletionSiqi Li, Changqing Zou, Yipeng Li, Xibin Zhao 等AAAI 2020 · 被引用 68 次
- VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene CompletionYiming Li, Zhiding Yu, Christopher B. Choy, Chaowei Xiao 等CVPR 2023
- Memory-Augmented Re-Completion for 3D Semantic Scene CompletionYu-Wen Tseng, Sheng-Ping Yang, Jhih-Ciang Wu, I-Bin Liao 等AAAI 2025 · 被引用 3 次
- NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates SpaceJiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai 等ICCV 2023 · 被引用 150 次
