DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
Bowen Yin, Jiao-Long Cao, Ming-Ming Cheng, Qibin Hou
摘要
Recent advances in scene understanding benefit a lot from depth maps because of the 3D geometry information, especially in complex conditions (e.g., low light and overexposed). Existing approaches encode depth maps along with RGB images and perform feature fusion between them to enable more robust predictions. Taking into account that depth can be regarded as a geometry supplement for RGB images, a straightforward question arises: Do we really need to explicitly encode depth information with neural networks as done for RGB images? Based on this insight, in this paper, we investigate a new way to learn RGBD feature representations and present DFormerv2, a strong RGBD encoder that explicitly uses depth maps as geometry priors rather than encoding depth information with neural networks. Our goal is to extract the geometry clues from the depth and spatial distances among all the image patch tokens, which will then be used as geometry priors to allocate attention weights in self-attention. Extensive experiments demonstrate that DFormerv2 exhibits exceptional performance in various RGBD semantic segmentation benchmarks. Code is available at: https://github.com/VCIP- RGBD/DFormer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic SegmentationBowen Yin, Jiao-Long Cao, Xuying Zhang, Yuming Chen 等NeurIPS 2025 · 被引用 8 次
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature AlignmentShi-Chen Zhang, Yunheng Li, Yu-Huan Wu, Qibin Hou 等ICCV 2025 · 被引用 8 次
- REL-SF4PASS: Panoramic Semantic Segmentation with REL Depth Representation and Spherical FusionXuewei Li, Xinghan Bao, Zhimin Chen, Xi LiCVPR 2026 · 被引用 3 次
- DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action RecognitionYuanjun Tan, Aoran Xiao, Liqian Deng, Zhigang TuCVPR 2026 · 被引用 1 次
- Unleashing Semantic and Geometric Priors for 3D Scene CompletionShiyuan Chen, Wei Sui, Bohao Zhang, Zeyd Boukhers 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper31
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu 等NeurIPS 2022 · 被引用 1,385 次
相关 Paper
- DFormer: Rethinking RGBD Representation Learning for Semantic SegmentationBowen Yin, Xuying Zhang, Zhong-Yu Li, Li Liu 等ICLR 2024 · 被引用 110 次
- Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal FusionPeng Sun, Wenhu Zhang, Huanyu Wang, Songyuan Li 等CVPR 2021
- Pri3D: Can 3D Priors Help 2D Representation Learning?Ji Hou, Saining Xie, Benjamin Graham, Angela Dai 等ICCV 2021 · 被引用 94 次
- SGFormer: Semantic-Geometry Fusion Transformer for Multi-modal 3D Panoptic SegmentationHongqi Yu, Sixian Chan, Xiaolong Zhou, Xiaoqin ZhangAAAI 2025 · 被引用 3 次
- ShapeConv: Shape-aware Convolutional Layer for Indoor RGB-D Semantic SegmentationJinming Cao, Hanchao Leng, Dani Lischinski, Danny Cohen-Or 等ICCV 2021 · 被引用 186 次
