TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding Network
Zhengyi Liu, Yuan Wang, Zhengzheng Tu, Yun Xiao, Bin Tang
Abstract
Salient object detection is the pixel-level dense prediction task which can highlight the prominent object in the scene. Recently U-Net framework is widely used, and continuous convolution and pooling operations generate multi-level features which are complementary with each other. In view of the more contribution of high-level features for the performance, we propose a triplet transformer embedding module to enhance them by learning long-range dependencies across layers. It is the first to use three transformer encoders with shared weights to enhance multi-level features. By further designing scale adjustment module to process the input, devising three-stream decoder to process the output and attaching depth features to color features for the multi-modal fusion, the proposed triplet transformer embedding network (TriTransNet) achieves the state-of-the-art performance in RGB-D salient object detection, and pushes the performance to a new level. Experimental results demonstrate the effectiveness of the proposed modules and the competition of TriTransNet. 1
• Computing methodologies → Interest point and salient region detections.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72e9e0fa-28c6-495f-9ba6-4456af23a9faCited by top-tier papers8
- Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object DetectionRunmin Cong, Hongyu Liu, Chen Zhang, Wei Zhang et al.ACM MM 2023 · 71 citations
- GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerDing Jia, Jianyuan Guo, Kai Han, Han Wu et al.ICML 2024 · 64 citations
- Object Segmentation by Mining Cross-Modal SemanticsZongwei Wu, Jingjing Wang, Zhuyun Zhou, Zhaochong An et al.ACM MM 2023 · 40 citations
- Weakly Supervised Video Salient Object Detection via Point SupervisionShuyong Gao, Haozhe Xing, Wei Zhang, Yan Wang et al.ACM MM 2022 · 39 citations
- Fantastic Animals and Where to Find Them: Segment Any Marine Animal with Dual SAMPingping Zhang, Tianyu Yan, Yang Liu, Huchuan LuCVPR 2024 · 32 citations
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo et al.NeurIPS 2021 · 2,148 citations
Related papers
- Visual Saliency TransformerNian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao et al.ICCV 2021 · 473 citations
- JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object DetectionKeren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun ZhaoCVPR 2020
- MMNet: Multi-Stage and Multi-Scale Fusion Network for RGB-D Salient Object DetectionGuibiao Liao, Wei Gao, Qiuping Jiang, Ronggang Wang et al.ACM MM 2020 · 53 citations
- Specificity-preserving RGB-D Saliency DetectionTao Zhou, Huazhu Fu, Geng Chen, Yi Zhou et al.ICCV 2021 · 210 citations
- Complementary Trilateral Decoder for Fast and Accurate Salient Object DetectionZhirui Zhao, Changqun Xia, Chenxi Xie, Jia LiACM MM 2021 · 119 citations
