MMNet: Multi-Stage and Multi-Scale Fusion Network for RGB-D Salient Object Detection
Guibiao Liao, Wei Gao, Qiuping Jiang, Ronggang Wang, Ge Li
Abstract
Most existing RGB-D salient object detection (SOD) methods directly extract and fuse raw features from RGB and depth backbones. Such methods can be easily restricted by low-quality depth maps and redundant cross-modal features. To effectively capture multi-scale cross-modal fusion features, this paper proposes a novel Multi-stage and Multi-Scale Fusion Network (MMNet), which consists of a cross-modal multi-stage fusion module (CMFM) and a bi-directional multi-scale decoder (BMD). Similar to the mechanism of visual color stage doctrine in human visual system, the proposed CMFM aims to explore the useful and important feature representations in feature response stage, and effectively integrate them into available cross-modal fusion features in adversarial combination stage. Moreover, the proposed BMD learns the combination of cross-modal fusion features from multiple levels to capture both local and global information of salient objects and further reasonably boost the performance of the proposed method. Comprehensive experiments demonstrate that the proposed method can achieve consistently superior performance over the other 14 state-of-the-art methods on six popular RGB-D datasets when evaluated by 8 different metrics.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 95eb33e9-456f-469e-9ad1-21e5d4c9a0bcCited by top-tier papers3
- Promoting Saliency From Depth: Deep Unsupervised RGB-D Saliency DetectionWei Ji, Jingjing Li, Qi Bi, Chuan Guo et al.ICLR 2022 · 46 citations
- End-to-End RGB-D Image Compression via Exploiting Channel-Modality RedundancyHuiming Zheng, Wei GaoAAAI 2024 · 15 citations
- SPC-GS: Gaussian Splatting with Semantic-Prompt Consistency for Indoor Open-World Free-view Synthesis from Sparse InputsGuibiao Liao, Qing Li, Zhenyu Bao, Guoping Qiu et al.CVPR 2025
Related papers
- Specificity-preserving RGB-D Saliency DetectionTao Zhou, Huazhu Fu, Geng Chen, Yi Zhou et al.ICCV 2021 · 210 citations
- Feature Reintegration over Differential Treatment: A Top-down and Adaptive Fusion Network for RGB-D Salient Object DetectionMiao Zhang, Yu Zhang, Yongri Piao, Beiqi Hu et al.ACM MM 2020 · 51 citations
- Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal FusionPeng Sun, Wenhu Zhang, Huanyu Wang, Songyuan Li et al.CVPR 2021
- RGB-D Salient Object Detection via 3D Convolutional Neural NetworksQian Chen, Ze Liu, Yi Zhang, Keren Fu et al.AAAI 2021 · 171 citations
- Cross-modality Discrepant Interaction Network for RGB-D Salient Object DetectionChen Zhang, Runmin Cong, Qinwei Lin, Lin Ma et al.ACM MM 2021 · 116 citations
