Unleashing Network Potentials for Semantic Scene Completion
Fengyun Wang, Qianru Sun, Dong Zhang, Jinhui Tang
Abstract
Semantic scene completion (SSC) aims to predict complete 3D voxel occupancy and semantics from a single-view RGB-D image, and recent SSC methods commonly adopt multi-modal inputs. However, our investigation reveals two limitations: ineffective feature learning from single modalities and overfitting to limited datasets. To address these issues, this paper proposes a novel SSC framework - Adversarial Modality Modulation Network (AMMNet) - with a fresh perspective of optimizing gradient updates. The proposed AMMNet introduces two core modules: a cross-modal modulation enabling the interdependence of gradient flows between modalities, and a customized adversarial training scheme leveraging dynamic gradient competition. Specifically, the cross-modal modulation adaptively re-calibrates the features to better excite representation potentials from each single modality. The adversarial training employs a minimax game of evolving gradients, with customized guidance to strengthen the generator's perception of visual fidelity from both geometric completeness and semantic correctness. Extensive experimental results demonstrate that AMMNet outperforms state-of-the-art SSC methods by a large margin, providing a promising direction for improving the effectiveness and generalization of SSC methods. Our code is available at this link.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7f8e3a88-4fdc-4397-b543-9caf655155fcCited by top-tier papers2
- Multi-modal Frequency Decomposition Network for Semantic Scene CompletionDie Zuo, Lubo Wang, Ruonan Liu, Qing Guo et al.CVPR 2026
- Point-based Instance Completion with Scene ConstraintsWesley Khademi, Fuxin LiICLR 2025
Builds on9
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Cascaded Context Pyramid for Full-Resolution 3D Semantic Scene CompletionPingping Zhang, Wei Liu, Yinjie Lei, Huchuan Lu et al.ICCV 2019 · 79 citations
- Not All Voxels Are Equal: Semantic Scene Completion from the Point-Voxel PerspectiveJiaxiang Tang, Xiaokang Chen, Jingbo Wang, Gang ZengAAAI 2022 · 37 citations
- FFNet: Frequency Fusion Network for Semantic Scene CompletionXuzhi Wang, Di Lin, Liang WanAAAI 2022 · 28 citations
Related papers
- Attention-Based Multi-Modal Fusion Network for Semantic Scene CompletionSiqi Li, Changqing Zou, Yipeng Li, Xibin Zhao et al.AAAI 2020 · 68 citations
- Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene CompletionYu Xue, Longjun Gao, Yuanqi Su, HaoAng Lu et al.CVPR 2026
- Anisotropic Convolutional Networks for 3D Semantic Scene CompletionJie Li, Kai Han, Peng Wang, Yu Liu et al.CVPR 2020
- MonoScene: Monocular 3D Semantic Scene CompletionAnh-Quan Cao, Raoul de CharetteCVPR 2022 · 251 citations
- RevealNet: Seeing Behind Objects in RGB-D ScansJi Hou, Angela Dai, Matthias NießnerCVPR 2020
