DFA3D: 3D Deformable Attention For 2D-to-3D Feature Lifting
Hongyang Li, Hao Zhang, Zhaoyang Zeng, Shilong Liu, Feng Li, Tianhe Ren, Lei Zhang
摘要
In this paper, we propose a new operator, called 3D DeFormable Attention (DFA3D), for 2D-to-3D feature lifting, which transforms multi-view 2D image features into a unified 3D space for 3D object detection. Existing feature lifting approaches, such as Lift-Splat-based and 2D attention-based, either use estimated depth to get pseudo LiDAR features and then splat them to a 3D space, which is a one-pass operation without feature refinement, or ignore depth and lift features by 2D attention mechanisms, which achieve finer semantics while suffering from a depth ambiguity problem. In contrast, our DFA3D-based method first leverages the estimated depth to expand each view's 2D feature map to 3D and then utilizes DFA3D to aggregate features from the expanded 3D feature maps. With the help of DFA3D, the depth ambiguity problem can be effectively alleviated from the root, and the lifted features can be progressively refined layer by layer, thanks to the Transformerlike architecture. In addition, we propose a mathematically equivalent implementation of DFA3D which can significantly improve its memory efficiency and computational speed. We integrate DFA3D into several methods that use 2D attention-based feature lifting with only a few modifications in code and evaluate on the nuScenes dataset. The experiment results show a consistent improvement of +1.41% mAP on average, and up to +15.1% mAP improvement when high-quality depth information is available, demonstrating the superiority, applicability, and huge potential of DFA3D. The code is available at https://github.com/IDEA-Research/3D-deformable-attention.git .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Context and Geometry Aware Voxel Transformer for Semantic Scene CompletionZhu Yu, Runmin Zhang, Jiacheng Ying, Junchen Yu 等NeurIPS 2024 · 被引用 73 次
- Multi-View Attentive Contextualization for Multi-View 3D Object DetectionXianpeng Liu, Ce Zheng, Ming Qian, Nan Xue 等CVPR 2024 · 被引用 5 次
- Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume ConstructionRunmin Zhang, Zhu Yu, Si-Yuan Cao, Lingyu Zhu 等ICCV 2025 · 被引用 3 次
- UniGS: Modeling Unitary 3D Gaussians for Novel View Synthesis from Sparse-View ImagesJiamin Wu, Kenkun Liu, Xiaoke Jiang, Yuan Yao 等ICCV 2025 · 被引用 3 次
- VisHall3D: Monocular Semantic Scene Completion from Reconstructing the Visible Regions to Hallucinating the Invisible RegionsHaoang Lu, Yuanqi Su, Xiaoning Zhang, Longjun Gao 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper13
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang 等AAAI 2023 · 被引用 954 次
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo 等CVPR 2022 · 被引用 879 次
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang 等ICLR 2023 · 被引用 753 次
相关 Paper
- Object as Query: Lifting any 2D Object Detector to 3D DetectionZitian Wang, Zehao Huang, Jiahui Fu, Naiyan Wang 等ICCV 2023 · 被引用 47 次
- FreqPDE: Rethinking Positional Depth Embedding for Multi-View 3D Object Detection TransformersHaisheng Su, Junjie Zhang, Feixiang Song, Sanping Zhou 等ICCV 2025
- LIFT: Learning 4D LiDAR Image Fusion Transformer for 3D Object DetectionYihan Zeng, Da Zhang, Chunwei Wang, Zhenwei Miao 等CVPR 2022 · 被引用 36 次
- Enhancing 3D Object Detection with 2D Detection-Guided Query AnchorsHaoxuanye Ji, Pengpeng Liang, Erkang ChengCVPR 2024
- 3DPPE: 3D Point Positional Encoding for Transformer-based Multi-Camera 3D Object DetectionChangyong Shu, Jiajun Deng, Fisher Yu, Yifan LiuICCV 2023 · 被引用 34 次
