MonoDGP: Monocular 3D Object Detection with Decoupled-Query and Geometry-Error Priors
Fanqi Pu, Yifan Wang, Jiru Deng, Wenming Yang
摘要
Perspective projection has been extensively utilized in monocular 3D object detection methods. It introduces geometric priors from 2D bounding boxes and 3D object dimensions to reduce the uncertainty of depth estimation. However, due to errors originating from the object's visual surface, the bounding box height often fails to represent the actual central height, which undermines the effectiveness of geometric depth. Direct prediction for the projected height unavoidably results in a loss of 2D priors, while multi-depth prediction with complex branches does not fully leverage geometric depth. This paper presents a Transformer-based monocular 3D object detection method called MonoDGP, which adopts perspective-invariant geometry errors to modify the projection formula. We also try to systematically discuss and explain the mechanisms and efficacy behind geometry errors, which serve as a simple but effective alternative to multi-depth prediction. Additionally, MonoDGP decouples the depth-guided decoder and constructs a 2D decoder only dependent on visual features, providing 2D priors and initializing object queries without the disturbance of 3D detection. To further optimize and finetune input tokens of the transformer decoder, we also introduce a Region Segmentation Head (RSH) that generates enhanced features and segment embeddings. Our monocular method demonstrates state-of-the-art performance on the KITTI benchmark without extra data. Code is available at https://github.com/PuFanqi23/MonoDGP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan, Xiaoming LiuCVPR 2026 · 被引用 13 次
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 被引用 5 次
- Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D DetectionXiang Li, Zhangchi Hu, Xu Xiao, Bin KongCVPR 2026 · 被引用 3 次
- MonoCLUE: Object-Aware Clustering Enhances Monocular 3D Object DetectionSunghun Yang, Minhyeok Lee, Jungho Lee, Sangyoun LeeAAAI 2026 · 被引用 2 次
- MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated LabelJunyoung Jung, Seokwon Kim, Jung Uk KimCVPR 2026 · 被引用 1 次
它引用的顶会 Paper23
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera 等ICCV 2019 · 被引用 504 次
- Geometry Uncertainty Projection Network for Monocular 3D Object DetectionYan Lu, Xinzhu Ma, Lei Yang, Tianzhu Zhang 等ICCV 2021 · 被引用 294 次
- Group DETR: Fast DETR Training with Group-Wise One-to-Many AssignmentQiang Chen, Xiaokang Chen, Jian Wang, Shan Zhang 等ICCV 2023 · 被引用 231 次
- How Do Neural Networks See Depth in Single Images?Tom van Dijk, Guido de CroonICCV 2019 · 被引用 210 次
相关 Paper
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo 等ICCV 2023 · 被引用 175 次
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 被引用 199 次
- Monocular 3D Object Detection with Decoupled Structured Polygon Estimation and Height-Guided Depth EstimationYingjie Cai, Buyu Li, Zeyu Jiao, Hongsheng Li 等AAAI 2020 · 被引用 100 次
- 3DPPE: 3D Point Positional Encoding for Transformer-based Multi-Camera 3D Object DetectionChangyong Shu, Jiajun Deng, Fisher Yu, Yifan LiuICCV 2023 · 被引用 34 次
- Geometry-based Distance Decomposition for Monocular 3D Object DetectionXuepeng Shi, Qi Ye, Xiaozhi Chen, Chuangrong Chen 等ICCV 2021 · 被引用 169 次
