Learning Selective Self-Mutual Attention for RGB-D Saliency Detection
Nian Liu, Ni Zhang, Junwei Han
Abstract
Saliency detection on RGB-D images is receiving more and more research interests recently. Previous models adopt the early fusion or the result fusion scheme to fuse the input RGB and depth data or their saliency maps, which incur the problem of distribution gap or information loss. Some other models use the feature fusion scheme but are limited by the linear feature fusion methods. In this paper, we propose to fuse attention learned in both modalities. Inspired by the Non-local model, we integrate the self-attention and each other's attention to propagate longrange contextual dependencies, thus incorporating multimodal information to learn attention and propagate contexts more accurately. Considering the reliability of the other modality's attention, we further propose a selection attention to weight the newly added attention term. We embed the proposed attention module in a two-stream C-NN for RGB-D saliency detection. Furthermore, we also propose a residual fusion module to fuse the depth decoder features into the RGB stream. Experimental results on seven benchmark datasets demonstrate the effectiveness of the proposed model components and our final saliency model. Our code and saliency maps are available at https://github.com/nnizhang/S2MA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 457fe441-5487-4574-817f-9bd1e4a3e38bCited by top-tier papers21
- Visual Saliency TransformerNian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao et al.ICCV 2021 · 473 citations
- Specificity-preserving RGB-D Saliency DetectionTao Zhou, Huazhu Fu, Geng Chen, Yi Zhou et al.ICCV 2021 · 210 citations
- TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding NetworkZhengyi Liu, Yuan Wang, Zhengzheng Tu, Yun Xiao et al.ACM MM 2021 · 175 citations
- Depth Quality-Inspired Feature Manipulation for Efficient RGB-D Salient Object DetectionWenbo Zhang, Ge-Peng Ji, Zhuo Wang, Keren Fu et al.ACM MM 2021 · 140 citations
- RGB-D Saliency Detection via Cascaded Mutual Information MinimizationJing Zhang, Deng-Ping Fan, Yuchao Dai, Xin Yu et al.ICCV 2021 · 122 citations
Builds on6
- CCNet: Criss-Cross Attention for Semantic SegmentationZilong Huang, Xinggang Wang, Lichao Huang, Chang Huang et al.ICCV 2019 · 2,972 citations
- Asymmetric Non-Local Neural Networks for Semantic SegmentationZhen Zhu, Mengdu Xu, Song Bai, Tengteng Huang et al.ICCV 2019 · 694 citations
- Expectation-Maximization Attention Networks for Semantic SegmentationXia Li, Zhisheng Zhong, Jianlong Wu, Yibo Yang et al.ICCV 2019 · 639 citations
- Depth-Induced Multi-Scale Recurrent Attention Network for Saliency DetectionYongri Piao, Wei Ji, Jingjing Li, Miao Zhang et al.ICCV 2019 · 450 citations
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox et al.ICCV 2019 · 157 citations
Related papers
- Select, Supplement and Focus for RGB-D Saliency DetectionMiao Zhang, Weisong Ren, Yongri Piao, Zhengkun Rong et al.CVPR 2020
- Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal FusionPeng Sun, Wenhu Zhang, Huanyu Wang, Songyuan Li et al.CVPR 2021
- JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object DetectionKeren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun ZhaoCVPR 2020
- MMNet: Multi-Stage and Multi-Scale Fusion Network for RGB-D Salient Object DetectionGuibiao Liao, Wei Gao, Qiuping Jiang, Ronggang Wang et al.ACM MM 2020 · 53 citations
- RGB-D Salient Object Detection via 3D Convolutional Neural NetworksQian Chen, Ze Liu, Yi Zhang, Keren Fu et al.AAAI 2021 · 171 citations
