Multi-modal Frequency Decomposition Network for Semantic Scene Completion
Die Zuo, Lubo Wang, Ruonan Liu, Qing Guo, Chong Wang, Dongdong Wu, Wei Feng, Kairui Yang, Di Lin
摘要
Based on an RGB-D image pair, semantic scene completion (SSC) provides a description for 3D scene understanding by predicting 3D semantic occupancy map. Recent methods extract RGB-D multi-modal features and fuse them in spatial domain, which disregards the misalignment caused by the imperfect raw multi-modal data and the multi-modal feature learning. Moreover, the operations of extracting high-level features they utilized tend to introduce feature smoothing and detail loss, exacerbating the above misalignment. To tackle these problems, this paper introduces MFDNet, a lightweight semantic scene completion network based on a multi-modal frequency decomposition strategy. By integrating frequency processing with limited layers of convolution and downsampling, MFDNet achieves a balance between modalities alignment and detail retainment. The network is equipped with Multi-modal Adaptive Frequency Fusion (MAFF) and Frequency Detail Compensation (FDC). MAFF models the intra-modal multi-bands dependencies and inter-modal relationships from a global perspective, enabling modality-specific calibration while facilitating the aligned fusion of multi-modal features. FDC excavates the high-frequency cues in shallow features to compensate for the missing local details of the fused feature and achieve fine-grained alignment for completion. MAFF and FDC formulate a global-to-local alignment and completion paradigm for multi-modal SSC. Extensive experiments demonstrate that MFDNet reduces parameters by 54.4% while achieving state-of-the-art performance on the NYUv2 and NYUCAD datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- NDC-Scene: Boost Monocular 3D Semantic Scene Completion in Normalized Device Coordinates SpaceJiawei Yao, Chuming Li, Keqiang Sun, Yingjie Cai 等ICCV 2023 · 被引用 150 次
- FilterNet: Harnessing Frequency Filters for Time Series ForecastingKun Yi, Jingru Fei, Qi Zhang, Hui He 等NeurIPS 2024 · 被引用 140 次
- XNet: Wavelet-Based Low and High Frequency Fusion Networks for Fully- and Semi-Supervised Semantic Segmentation of Biomedical ImagesYanfeng Zhou, Jiaxing Huang, Chenlong Wang, Le Song 等ICCV 2023 · 被引用 89 次
- FreGS: 3D Gaussian Splatting with Progressive Frequency RegularizationJiahui Zhang, Fangneng Zhan, Muyu Xu, Shijian Lu 等CVPR 2024 · 被引用 61 次
- FFNet: Frequency Fusion Network for Semantic Scene CompletionXuzhi Wang, Di Lin, Liang WanAAAI 2022 · 被引用 28 次
相关 Paper
- Attention-Based Multi-Modal Fusion Network for Semantic Scene CompletionSiqi Li, Changqing Zou, Yipeng Li, Xibin Zhao 等AAAI 2020 · 被引用 68 次
- Unleashing Network Potentials for Semantic Scene CompletionFengyun Wang, Qianru Sun, Dong Zhang, Jinhui TangCVPR 2024 · 被引用 3 次
- Completion as Enhancement: A Degradation-Aware Selective Image Guided Network for Depth CompletionZhiqiang Yan, Zhengxue Wang, Kun Wang, Jun Li 等CVPR 2025
- Semantic Scene Completion with Cleaner SelfFengyun Wang, Dong Zhang, Hanwang Zhang, Jinhui Tang 等CVPR 2023
- SDNet: LiDAR Semantic Scene Completion with Sparse-Dense Fusion and Input-Aware Label RefinementTingming Bai, Zhiyu Xiang, Peng Xu, Tianyu Pu 等AAAI 2026
