Joint Semantic Mining for Weakly Supervised RGB-D Salient Object Detection
Jingjing Li, Wei Ji, Qi Bi, Cheng Yan, Miao Zhang, Yongri Piao, Huchuan Lu, Li Cheng
摘要
Training saliency detection models with weak supervisions, e.g., image-level tags or captions, is appealing as it removes the costly demand of per-pixel annotations. Despite the rapid progress of RGB-D saliency detection in fully-supervised setting, it however remains an unexplored territory when only weak supervision signals are available. This paper is set to tackle the problem of weakly-supervised RGB-D salient object detection. The key insight in this effort is the idea of maintaining perpixel pseudo-labels with iterative refinements by reconciling the multimodal input signals in our joint semantic mining (JSM). Considering the large variations in the raw depth map and the lack of explicit pixel-level supervisions, we propose spatial semantic modeling (SSM) to capture saliency-specific depth cues from the raw depth and produce depth-refined pseudo-labels. Moreover, tags and captions are incorporated via a fill-in-the-blank training in our textual semantic modeling (TSM) to estimate the confidences of competing pseudo-labels. At test time, our model involves only a light-weight sub-network of the training pipeline, i.e., it requires only an RGB image as input, thus allowing efficient inference. Extensive evaluations demonstrate the effectiveness of our approach under the weakly-supervised setting. Importantly, our method could also be adapted to work in both fully-supervised and unsupervised paradigms. In each of these scenarios, superior performance has been attained by our approach with comparing to the state-of-the-art dedicated methods. As a by-product, a CapS dataset is constructed by augmenting existing benchmark training set with additional image tags and captions. Code and dataset are available at https://github.com/jiwei0921/JSM . RGB Image Depth Map Initial Pseudo-label Updated Label 1 Updated Label 3 Updated Label 7 Caption: A black and white cat is laying on a gray couch.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Zoom In and Out: A Mixed-scale Triplet Network for Camouflaged Object DetectionYouwei Pang, Xiaoqi Zhao, Tian-Zhu Xiang, Lihe Zhang 等CVPR 2022 · 被引用 417 次
- Self-Supervised Pretraining for RGB-D Salient Object DetectionXiaoqi Zhao, Youwei Pang, Lihe Zhang, Huchuan Lu 等AAAI 2022 · 被引用 78 次
- Learning Content-Enhanced Mask Transformer for Domain Generalized Urban-Scene SegmentationQi Bi, Shaodi You, Theo GeversAAAI 2024 · 被引用 77 次
- Learning Frequency-Adapted Vision Foundation Model for Domain Generalized Semantic SegmentationQi Bi, Jingjun Yi, Hao Zheng, Haolan Zhan 等NeurIPS 2024 · 被引用 62 次
- Promoting Saliency From Depth: Deep Unsupervised RGB-D Saliency DetectionWei Ji, Jingjing Li, Qi Bi, Chuan Guo 等ICLR 2022 · 被引用 46 次
它引用的顶会 Paper22
- Depth-Induced Multi-Scale Recurrent Attention Network for Saliency DetectionYongri Piao, Wei Ji, Jingjing Li, Miao Zhang 等ICCV 2019 · 被引用 450 次
- Specificity-preserving RGB-D Saliency DetectionTao Zhou, Huazhu Fu, Geng Chen, Yi Zhou 等ICCV 2021 · 被引用 210 次
- Structure-Consistent Weakly Supervised Salient Object Detection with Local Saliency CoherenceSiyue Yu, Bingfeng Zhang, Jimin Xiao, Eng Gee LimAAAI 2021 · 被引用 162 次
- Learning Generative Vision Transformer with Energy-Based Latent Space for Saliency PredictionJing Zhang, Jianwen Xie, Nick Barnes, Ping LiNeurIPS 2021 · 被引用 117 次
- Dynamic Context-Sensitive Filtering Network for Video Salient Object DetectionMiao Zhang, Jie Liu, Yifei Wang, Yongri Piao 等ICCV 2021 · 被引用 112 次
相关 Paper
- Railroad Is Not a Train: Saliency As Pseudo-Pixel Supervision for Weakly Supervised Semantic SegmentationSeungho Lee, Minhyun Lee, Jongwuk Lee, Hyunjung ShimCVPR 2021
- JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object DetectionKeren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun ZhaoCVPR 2020
- Weakly-Supervised Salient Object Detection via Scribble AnnotationsJing Zhang, Xin Yu, Aixuan Li, Peipei Song 等CVPR 2020
- Weakly-Supervised Salient Object Detection Using Point SupervisonShuyong Gao, Wei Zhang, Yan Wang, Qianyu Guo 等AAAI 2022 · 被引用 77 次
- Weakly Supervised Video Salient Object DetectionWangbo Zhao, Jing Zhang, Long Li, Nick Barnes 等CVPR 2021
