TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding Network
Zhengyi Liu, Yuan Wang, Zhengzheng Tu, Yun Xiao, Bin Tang
摘要
Salient object detection is the pixel-level dense prediction task which can highlight the prominent object in the scene. Recently U-Net framework is widely used, and continuous convolution and pooling operations generate multi-level features which are complementary with each other. In view of the more contribution of high-level features for the performance, we propose a triplet transformer embedding module to enhance them by learning long-range dependencies across layers. It is the first to use three transformer encoders with shared weights to enhance multi-level features. By further designing scale adjustment module to process the input, devising three-stream decoder to process the output and attaching depth features to color features for the multi-modal fusion, the proposed triplet transformer embedding network (TriTransNet) achieves the state-of-the-art performance in RGB-D salient object detection, and pushes the performance to a new level. Experimental results demonstrate the effectiveness of the proposed modules and the competition of TriTransNet. 1
• Computing methodologies → Interest point and salient region detections.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Point-aware Interaction and CNN-induced Refinement Network for RGB-D Salient Object DetectionRunmin Cong, Hongyu Liu, Chen Zhang, Wei Zhang 等ACM MM 2023 · 被引用 71 次
- GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerDing Jia, Jianyuan Guo, Kai Han, Han Wu 等ICML 2024 · 被引用 64 次
- Object Segmentation by Mining Cross-Modal SemanticsZongwei Wu, Jingjing Wang, Zhuyun Zhou, Zhaochong An 等ACM MM 2023 · 被引用 40 次
- Weakly Supervised Video Salient Object Detection via Point SupervisionShuyong Gao, Haozhe Xing, Wei Zhang, Yan Wang 等ACM MM 2022 · 被引用 39 次
- Fantastic Animals and Where to Find Them: Segment Any Marine Animal with Dual SAMPingping Zhang, Tianyu Yan, Yang Liu, Huchuan LuCVPR 2024 · 被引用 32 次
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu 等ICCV 2021 · 被引用 2,397 次
- Transformer in TransformerKai Han, An Xiao, Enhua Wu, Jianyuan Guo 等NeurIPS 2021 · 被引用 2,148 次
相关 Paper
- Visual Saliency TransformerNian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao 等ICCV 2021 · 被引用 473 次
- JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object DetectionKeren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun ZhaoCVPR 2020
- MMNet: Multi-Stage and Multi-Scale Fusion Network for RGB-D Salient Object DetectionGuibiao Liao, Wei Gao, Qiuping Jiang, Ronggang Wang 等ACM MM 2020 · 被引用 53 次
- Specificity-preserving RGB-D Saliency DetectionTao Zhou, Huazhu Fu, Geng Chen, Yi Zhou 等ICCV 2021 · 被引用 210 次
- Complementary Trilateral Decoder for Fast and Accurate Salient Object DetectionZhirui Zhao, Changqun Xia, Chenxi Xie, Jia LiACM MM 2021 · 被引用 119 次
