Joint Semantic Mining for Weakly Supervised RGB-D Salient Object Detection
Jingjing Li, Wei Ji, Qi Bi, Cheng Yan, Miao Zhang, Yongri Piao, Huchuan Lu, Li Cheng
Abstract
Training saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection in fully-supervised setting, it however remains an unexplored territory when only weak supervision signals are available. This paper is set to tackle the problem of weakly-supervised RGB-D salient object detection. The key insight in this effort is the idea of maintaining perpixel pseudo-labels with iterative refinements by reconciling the multimodal input signals in our joint semantic mining (JSM). Considering the large variations in the raw depth map and the lack of explicit pixel-level supervisions, we propose spatial semantic modeling (SSM) to capture saliency-specific depth cues from the raw depth and produce depth-refined pseudo-labels. Moreover, tags and captions are incorporated via a fill-in-the-blank training in our textual semantic modeling (TSM) to estimate the confidences of competing pseudo-labels. At test time, our model involves only a light-weight sub-network of the training pipeline, i.e., it requires only an RGB image as input, thus allowing efficient inference. Extensive evaluations demonstrate the effectiveness of our approach under the weakly-supervised setting. Importantly, our method could also be adapted to work in both fully-supervised and unsupervised paradigms. In each of these scenarios, superior performance has been attained by our approach with comparing to the state-of-the-art dedicated methods. As a by-product, a CapS dataset is constructed by augmenting existing benchmark training set with additional image tags and captions. Code and dataset are available at https://github.com/jiwei0921/JSM . RGB Image Depth Map Initial Pseudo-label Updated Label 1 Updated Label 3 Updated Label 7 Caption: A black and white cat is laying on a gray couch.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 09323665-fd0f-4abc-9e33-a34c9799cc91Cited by top-tier papers13
- Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object DetectionYouwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang et al.CVPR 2022 · 417 citations
- Self-Supervised Pretraining for RGB-D Salient Object DetectionXiaoqi Zhao, Youwei Pang, Lihe Zhang, Huchuan Lu et al.AAAI 2022 · 78 citations
- Learning Content-Enhanced Mask Transformer for Domain Generalized Urban-Scene SegmentationQi Bi, Shaodi You, Theo GeversAAAI 2024 · 77 citations
- Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic SegmentationQi Bi, Jingjun Yi, Hao Zheng, Haolan Zhan et al.NeurIPS 2024 · 62 citations
- Promoting Saliency From Depth: Deep Unsupervised RGB-D Saliency DetectionWei Ji, Jingjing Li, Qi Bi, Chuan Guo et al.ICLR 2022 · 46 citations
Builds on22
- Depth-Induced Multi-Scale Recurrent Attention Network for Saliency DetectionYongri Piao, Wei Ji, Jingjing Li, Miao Zhang et al.ICCV 2019 · 450 citations
- Specificity-preserving RGB-D Saliency DetectionTao Zhou, Huazhu Fu, Geng Chen, Yi Zhou et al.ICCV 2021 · 210 citations
- Structure-Consistent Weakly Supervised Salient Object Detection with Local Saliency CoherenceSiyue Yu, Bingfeng Zhang, Jimin Xiao, Eng Gee LimAAAI 2021 · 162 citations
- Learning Generative Vision Transformer with Energy-Based Latent Space for Saliency PredictionJing Zhang, Jianwen Xie, Nick Barnes, Ping LiNeurIPS 2021 · 117 citations
- Dynamic Context-Sensitive Filtering Network for Video Salient Object DetectionMiao Zhang, Jie Liu, Yifei Wang, Yongri Piao et al.ICCV 2021 · 112 citations
Related papers
- Railroad Is Not a Train: Saliency As Pseudo-Pixel Supervision for Weakly Supervised Semantic SegmentationSeungho Lee, Minhyun Lee, Jongwuk Lee, Hyunjung ShimCVPR 2021
- JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object DetectionKeren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun ZhaoCVPR 2020
- Weakly-Supervised Salient Object Detection via Scribble AnnotationsJing Zhang, Xin Yu, Aixuan Li, Peipei Song et al.CVPR 2020
- Weakly-Supervised Salient Object Detection Using Point SupervisonShuyong Gao, Wei Zhang, Yan Wang, Qianyu Guo et al.AAAI 2022 · 77 citations
- Weakly Supervised Video Salient Object DetectionWangbo Zhao, Jing Zhang, Long Li, Nick Barnes et al.CVPR 2021
