Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object Detection
Runmin Cong, Hongyu Liu, Chen Zhang, Wei Zhang, Feng Zheng, Ran Song, Sam Kwong
Abstract
By integrating complementary information from RGB image and depth map, the ability of salient object detection (SOD) for complex and challenging scenes can be improved. In recent years, the important role of Convolutional Neural Networks (CNNs) in feature extraction and cross-modality interaction has been fully explored, but it is still insufficient in modeling global long-range dependencies of self-modality and cross-modality. To this end, we introduce CNNs-assisted Transformer architecture and propose a novel RGB-D SOD network with Point-aware Interaction and CNN-induced Refinement (PICR-Net). On the one hand, considering the prior correlation between RGB modality and depth modality, an attentiontriggered cross-modality point-aware interaction (CmPI) module is designed to explore the feature interaction of different modalities with positional constraints. On the other hand, in order to alleviate the block effect and detail destruction problems brought by the Transformer naturally, we design a CNN-induced refinement (CNNR) unit for content refinement and supplementation. Extensive experiments on five RGB-D SOD datasets show that the proposed network achieves competitive results in both quantitative and qualitative comparisons. Our code is publicly available at: https:// github.com/ rmcong/ PICR-Net_ACMMM23.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fac65224-b879-4e7e-bd05-19815cc83e62Cited by top-tier papers4
- ESNet: Evolution and Succession Network for High-Resolution Salient Object DetectionHongyu Liu, Runmin Cong, Hua Li, Qianqian Xu et al.ICML 2024 · 7 citations
- LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object DetectionLanhu Wu, Zilin Gao, Hao Fei, Mong-Li Lee et al.ACM MM 2025 · 3 citations
- M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object DetectionJiyuan Liu, Jia Lin, Xiaofei Zhou, Runmin Cong et al.CVPR 2026
- SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object DetectionJia Lin, Xiaofei Zhou, Jiyuan Liu, Runmin Cong et al.AAAI 2026
Builds on16
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- PANet: Few-Shot Image Semantic Segmentation With Prototype AlignmentKaixin Wang, Jun Hao Liew, Yingtian Zou, Daquan Zhou et al.ICCV 2019 · 1,404 citations
- Global Context-Aware Progressive Aggregation Network for Salient Object DetectionZuyao Chen, Qianqian Xu, Runmin Cong, Qingming HuangAAAI 2020 · 481 citations
- Visual Saliency TransformerNian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao et al.ICCV 2021 · 473 citations
Related papers
- Cross-modality Discrepant Interaction Network for RGB-D Salient Object DetectionChen Zhang, Runmin Cong, Qinwei Lin, Lin Ma et al.ACM MM 2021 · 116 citations
- MMNet: Multi-Stage and Multi-Scale Fusion Network for RGB-D Salient Object DetectionGuibiao Liao, Wei Gao, Qiuping Jiang, Ronggang Wang et al.ACM MM 2020 · 53 citations
- Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal FusionPeng Sun, Wenhu Zhang, Huanyu Wang, Songyuan Li et al.CVPR 2021
- RGB-D Salient Object Detection via 3D Convolutional Neural NetworksQian Chen, Ze Liu, Yi Zhang, Keren Fu et al.AAAI 2021 · 171 citations
- Is Depth Really Necessary for Salient Object Detection?Jiawei Zhao, Yifan Zhao, Jia Li, Xiaowu ChenACM MM 2020 · 71 citations
