V-DETR: DETR with Vertex Relative Position Encoding for 3D Object Detection
Yichao Shen, Zigang Geng, Yuhui Yuan, Yutong Lin, Ze Liu, Chunyu Wang, Han Hu, Nanning Zheng, Baining Guo
Abstract
We introduce a highly performant 3D object detector for point clouds using the DETR framework. The prior attempts all end up with suboptimal results because they fail to learn accurate inductive biases from the limited scale of training data. In particular, the queries often attend to points that are far away from the target objects, violating the locality principle in object detection. To address the limitation, we introduce a novel 3D Vertex Relative Position Encoding (3DV-RPE) method which computes position encoding for each point based on its relative position to the 3D boxes predicted by the queries in each decoder layer, thus providing clear information to guide the model to focus on points near the objects, in accordance with the principle of locality. In addition, we systematically improve the pipeline from various aspects such as data normalization based on our understanding of the task. We show exceptional results on the challenging ScanNetV2 benchmark, achieving significant improvements over the previous 3DETR in / from 65.0%/47.0% to 77.8%/66.0%, respectively. In addition, our method sets a new record on ScanNetV2 and SUN RGB-D datasets.Code will be released at http://github.com/yichaoshen-MS/V-DETR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdc8ea8f-9b86-4abf-9e50-c9aa61d5a5b9Cited by top-tier papers14
- SpatialLM: Training Large Language Models for Structured Indoor ModelingYongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng et al.NeurIPS 2025 · 89 citations
- Training an Open-Vocabulary Monocular 3D Detection Model without 3D DataRui Huang, Henry Zheng, Yan Wang, Zhuofan Xia et al.NeurIPS 2024 · 26 citations
- UniDet3D: Multi-dataset Indoor 3D Object DetectionMaksim Kolodiazhnyi, Anna Vorontsova, Matvey Skripkin, Danila Rukhovich et al.AAAI 2025 · 7 citations
- Zoo3D: Zero-Shot 3D Object Detection at Scene LevelAndrey Lemeshko, Bulat Gabdullin, Nikita Drozdov, Anton Konushin et al.CVPR 2026 · 5 citations
- SA3DIP: Segment Any 3D Instance with Potential 3D PriorsXi Yang, Xu Gu, Xingyilang Yin, Xinbo GaoNeurIPS 2024 · 3 citations
Builds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
Related papers
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 602 citations
- Back-Tracing Representative Points for Voting-Based 3D Object Detection in Point CloudsBowen Cheng, Lu Sheng, Shaoshuai Shi, Ming Yang et al.CVPR 2021
- Uni3DETR: Unified 3D Detection TransformerZhenyu Wang, Ya-Li Li, Xi Chen, Hengshuang Zhao et al.NeurIPS 2023 · 65 citations
- RBGNet: Ray-based Grouping for 3D Object DetectionHaiyang Wang, Shaoshuai Shi, Ze Yang, Rongyao Fang et al.CVPR 2022 · 63 citations
- DisARM: Displacement Aware Relation Module for 3D DetectionYao Duan, Chenyang Zhu, Yuqing Lan, Renjiao Yi et al.CVPR 2022 · 22 citations
