Self-Supervised Pretraining for RGB-D Salient Object Detection
Xiaoqi Zhao, Youwei Pang, Lihe Zhang, Huchuan Lu, Xiang Ruan
Abstract
Existing CNNs-Based RGB-D salient object detection (SOD) networks are all required to be pretrained on the ImageNet to learn the hierarchy features which helps provide a good initialization. However, the collection and annotation of largescale datasets are time-consuming and expensive. In this paper, we utilize self-supervised representation learning (SSL) to design two pretext tasks: the cross-modal auto-encoder and the depth-contour estimation. Our pretext tasks require only a few and unlabeled RGB-D datasets to perform pretraining, which makes the network capture rich semantic contexts and reduce the gap between two modalities, thereby providing an effective initialization for the downstream task. In addition, for the inherent problem of cross-modal fusion in RGB-D SOD, we propose a consistency-difference aggregation (CDA) module that splits a single feature fusion into multi-path fusion to achieve an adequate perception of consistent and differential information. The CDA module is general and suitable for cross-modal and cross-level feature fusion. Extensive experiments on six benchmark datasets show that our self-supervised pretrained model performs favorably against most state-of-the-art methods pretrained on Im-ageNet. The source code will be publicly available at https: //github.com/Xiaoqi-Zhao-DLUT/SSLSOD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc2b200e-9b27-41de-960b-ebb81e417b15Cited by top-tier papers7
- Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object DetectionYouwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang et al.CVPR 2022 · 417 citations
- Promoting Saliency From Depth: Deep Unsupervised RGB-D Saliency DetectionWei Ji, Jingjing Li, Qi Bi, Chuan Guo et al.ICLR 2022 · 46 citations
- Spider: A Unified Framework for Context-dependent Concept SegmentationXiaoqi Zhao, Youwei Pang, Wei Ji, Baicheng Sheng et al.ICML 2024 · 21 citations
- Multi-View Aggregation Network for Dichotomous Image SegmentationQian Yu, Xiaoqi Zhao, Youwei Pang, Lihe Zhang et al.CVPR 2024 · 13 citations
- Self-supervised Pre-training for Mirror DetectionJiaying Lin, Rynson W. H. LauICCV 2023 · 9 citations
Builds on13
- Depth-Induced Multi-Scale Recurrent Attention Network for Saliency DetectionYongri Piao, Wei Ji, Jingjing Li, Miao Zhang et al.ICCV 2019 · 450 citations
- Scaling and Benchmarking Self-Supervised Visual Representation LearningPriya Goyal, Dhruv Mahajan, Abhinav Gupta, Ishan MisraICCV 2019 · 429 citations
- Depth Quality-Inspired Feature Manipulation for Efficient RGB-D Salient Object DetectionWenbo Zhang, Ge-Peng Ji, Zhuo Wang, Keren Fu et al.ACM MM 2021 · 140 citations
- Joint Semantic Mining for Weakly Supervised RGB-D Salient Object DetectionJingjing Li, Wei Ji, Qi Bi, Cheng Yan et al.NeurIPS 2021 · 56 citations
- You Only Infer Once: Cross-Modal Meta-Transfer for Referring Video Object SegmentationDezhuang Li, Ruoqi Li, Lijun Wang, Yifan Wang et al.AAAI 2022 · 54 citations
Related papers
- RGB-D Salient Object Detection via 3D Convolutional Neural NetworksQian Chen, Ze Liu, Yi Zhang, Keren Fu et al.AAAI 2021 · 171 citations
- Cross-modality Discrepant Interaction Network for RGB-D Salient Object DetectionChen Zhang, Runmin Cong, Qinwei Lin, Lin Ma et al.ACM MM 2021 · 116 citations
- CoMAE: Single Model Hybrid Pre-training on Small-Scale RGB-D DatasetsJiange Yang, Sheng Guo, Gangshan Wu, Limin WangAAAI 2023 · 12 citations
- Specificity-preserving RGB-D Saliency DetectionTao Zhou, Huazhu Fu, Geng Chen, Yi Zhou et al.ICCV 2021 · 210 citations
- Self-Supervised Pretraining of 3D Features on any Point-CloudZaiwei Zhang, Rohit Girdhar, Armand Joulin, Ishan MisraICCV 2021 · 333 citations
