DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
Bowen Yin, Xuying Zhang, Zhong-Yu Li, Li Liu, Ming-Ming Cheng, Qibin Hou
Abstract
We present DFormer, a novel RGB-D pretraining framework to learn transferable representations for RGB-D segmentation tasks. DFormer has two new key innovations: 1) Unlike previous works that encode RGB-D information with RGB pretrained backbone, we pretrain the backbone using image-depth pairs from ImageNet-1K, and hence the DFormer is endowed with the capacity to encode RGB-D representations; 2) DFormer comprises a sequence of RGB-D blocks, which are tailored for encoding both RGB and depth information through a novel building block design. DFormer avoids the mismatched encoding of the 3D geometry relationships in depth maps by RGB pretrained backbones, which widely lies in existing methods but has not been resolved. We finetune the pretrained DFormer on two popular RGB-D tasks, i.e., RGB-D semantic segmentation and RGB-D salient object detection, with a lightweight decoder head. Experimental results show that our DFormer achieves new state-of-the-art performance on these two tasks with less than half of the computational cost of the current best methods on two RGB-D semantic segmentation datasets and five RGB-D salient object detection datasets. Our code is available at: https://github.com/VCIP-RGBD/DFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision TransformerDing Jia, Jianyuan Guo, Kai Han, Han Wu et al.ICML 2024 · 64 citations
- Strip R-CNN: Large Strip Convolution for Remote Sensing Object DetectionXinbin Yuan, Zhaohui Zheng, Yuxuan Li, Xialei Liu et al.AAAI 2026 · 32 citations
- StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic SegmentationBingyu Li, Da Zhang, Zhiyuan Zhao, Junyu Gao et al.ACM MM 2025 · 15 citations
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic SegmentationBowen Yin, Jiao-Long Cao, Xuying Zhang, Yuming Chen et al.NeurIPS 2025 · 8 citations
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature AlignmentShi-Chen Zhang, Yunheng Li, Yu-Huan Wu, Qibin Hou et al.ICCV 2025 · 8 citations
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and TextHassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang et al.NeurIPS 2021 · 782 citations
Related papers
- DFormerv2: Geometry Self-Attention for RGBD Semantic SegmentationBowen Yin, Jiao-Long Cao, Ming-Ming Cheng, Qibin HouCVPR 2025
- JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object DetectionKeren Fu, Deng-Ping Fan, Ge-Peng Ji, Qijun ZhaoCVPR 2020
- Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal FusionPeng Sun, Wenhu Zhang, Huanyu Wang, Songyuan Li et al.CVPR 2021
- Pri3D: Can 3D Priors Help 2D Representation Learning?Ji Hou, Saining Xie, Benjamin Graham, Angela Dai et al.ICCV 2021 · 94 citations
- Cross-modality Discrepant Interaction Network for RGB-D Salient Object DetectionChen Zhang, Runmin Cong, Qinwei Lin, Lin Ma et al.ACM MM 2021 · 116 citations
