Attention-Based Multi-Modal Fusion Network for Semantic Scene Completion
Siqi Li, Changqing Zou, Yipeng Li, Xibin Zhao, Yue Gao
Abstract
This paper presents an end-to-end 3D convolutional network named attention-based multi-modal fusion network (AMFNet) for the semantic scene completion (SSC) task of inferring the occupancy and semantic labels of a volumetric 3D scene from single-view RGB-D images. Compared with previous methods which use only the semantic features extracted from RGB-D images, the proposed AMFNet learns to perform effective 3D scene completion and semantic segmentation simultaneously via leveraging the experience of inferring 2D semantic segmentation from RGB-D images as well as the reliable depth cues in spatial dimension. It is achieved by employing a multi-modal fusion architecture boosted from 2D semantic segmentation and a 3D semantic completion network empowered by residual attention blocks. We validate our method on both the synthetic SUNCG-RGBD dataset and the real NYUv2 dataset and the results show that our method respectively achieves the gains of 2.5% and 2.6% on the synthetic SUNCG-RGBD dataset and the real NYUv2 dataset against the state-of-the-art method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 992cbbcf-1649-4d6e-8ec6-3870acd36077Cited by top-tier papers19
- OpenOccupancy: A Large Scale Benchmark for Surrounding Semantic Occupancy PerceptionXiaofeng Wang, Zheng Zhu, Wenbo Xu, Yunpeng Zhang et al.ICCV 2023 · 270 citations
- MonoScene: Monocular 3D Semantic Scene CompletionAnh-Quan Cao, Raoul de CharetteCVPR 2022 · 251 citations
- ShapeConv: Shape-aware Convolutional Layer for Indoor RGB-D Semantic SegmentationJinming Cao, Hanchao Leng, Dani Lischinski, Danny Cohen-Or et al.ICCV 2021 · 186 citations
- Learning Multi-View Aggregation In the Wild for Large-Scale 3D Semantic SegmentationDamien Robert, Bruno Vallet, Loïc LandrieuCVPR 2022 · 84 citations
- Not All Voxels Are Equal: Semantic Scene Completion from the Point-Voxel PerspectiveJiaxiang Tang, Xiaokang Chen, Jingbo Wang, Gang ZengAAAI 2022 · 37 citations
Related papers
- Multi-modal Frequency Decomposition Network for Semantic Scene CompletionDie Zuo, Lubo Wang, Ruonan Liu, Qing Guo et al.CVPR 2026
- Unleashing Network Potentials for Semantic Scene CompletionFengyun Wang, Qianru Sun, Dong Zhang, Jinhui TangCVPR 2024 · 3 citations
- FFNet: Frequency Fusion Network for Semantic Scene CompletionXuzhi Wang, Di Lin, Liang WanAAAI 2022 · 28 citations
- Cascaded Context Pyramid for Full-Resolution 3D Semantic Scene CompletionPingping Zhang, Wei Liu, Yinjie Lei, Huchuan Lu et al.ICCV 2019 · 79 citations
- Anisotropic Convolutional Networks for 3D Semantic Scene CompletionJie Li, Kai Han, Peng Wang, Yu Liu et al.CVPR 2020
