Multi-modal Frequency Decomposition Network for Semantic Scene Completion
Die Zuo, Lubo Wang, Ruonan Liu, Qing Guo, Chong Wang, Dongdong Wu, Wei Feng, Kairui Yang, Di Lin
Abstract
Based on an RGB-D image pair, semantic scene completion (SSC) provides a description for 3D scene understanding by predicting 3D semantic occupancy map. Recent methods extract RGB-D multi-modal features and fuse them in spatial domain, which disregards the misalignment caused by the imperfect raw multi-modal data and the multi-modal feature learning. Moreover, the operations of extracting high-level features they utilized tend to introduce feature smoothing and detail loss, exacerbating the above misalignment. To tackle these problems, this paper introduces MFDNet, a lightweight semantic scene completion network based on a multi-modal frequency decomposition strategy. By integrating frequency processing with limited layers of convolution and downsampling, MFDNet achieves a balance between modalities alignment and detail retainment. The network is equipped with Multi-modal Adaptive Frequency Fusion (MAFF) and Frequency Detail Compensation (FDC). MAFF models the intra-modal multi-bands dependencies and inter-modal relationships from a global perspective, enabling modality-specific calibration while facilitating the aligned fusion of multi-modal features. FDC excavates the high-frequency cues in shallow features to compensate for the missing local details of the fused feature and achieve fine-grained alignment for completion. MAFF and FDC formulate a global-to-local alignment and completion paradigm for multi-modal SSC. Extensive experiments demonstrate that MFDNet reduces parameters by 54.4% while achieving state-of-the-art performance on the NYUv2 and NYUCAD datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates SpaceJiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai et al.ICCV 2023 · 150 citations
- FilterNet: Harnessing Frequency Filters for Time Series ForecastingKun Yi, Jingru Fei, Qi Zhang, Hui He et al.NeurIPS 2024 · 140 citations
- XNet: Wavelet-Based Low and High Frequency Fusion Networks for Fully- and Semi-Supervised Semantic Segmentation of Biomedical ImagesYanfeng Zhou, Jiaxing Huang, Chenlong Wang, Le Song et al.ICCV 2023 · 89 citations
- FreGS: 3D Gaussian Splatting with Progressive Frequency RegularizationJiahui Zhang, Fangneng Zhan, Muyu Xu, Shijian Lu et al.CVPR 2024 · 61 citations
- FFNet: Frequency Fusion Network for Semantic Scene CompletionXuzhi Wang, Di Lin, Liang WanAAAI 2022 · 28 citations
Related papers
- Attention-Based Multi-Modal Fusion Network for Semantic Scene CompletionSiqi Li, Changqing Zou, Yipeng Li, Xibin Zhao et al.AAAI 2020 · 68 citations
- Unleashing Network Potentials for Semantic Scene CompletionFengyun Wang, Qianru Sun, Dong Zhang, Jinhui TangCVPR 2024 · 3 citations
- Completion as Enhancement: A Degradation-Aware Selective Image Guided Network for Depth CompletionZhiqiang Yan, Zhengxue Wang, Kun Wang, Jun Li et al.CVPR 2025
- Semantic Scene Completion with Cleaner SelfFengyun Wang, Dong Zhang, Hanwang Zhang, Jinhui Tang et al.CVPR 2023
- SDNet: LiDAR Semantic Scene Completion with Sparse-Dense Fusion and Input-Aware Label RefinementTingming Bai, Zhiyu Xiang, Peng Xu, Tianyu Pu et al.AAAI 2026
