Geometry Uncertainty Projection Network for Monocular 3D Object Detection
Yan Lu, Xinzhu Ma, Lei Yang, Tianzhu Zhang, Yating Liu, Qi Chu, Junjie Yan, Wanli Ouyang
Abstract
Geometry Projection is a powerful depth estimation method in monocular 3D object detection. It estimates depth dependent on heights, which introduces mathematical priors into the deep model. But projection process also introduces the error amplification problem, in which the error of the estimated height will be amplified and reflected greatly at the output depth. This property leads to uncontrollable depth inferences and also damages the training efficiency. In this paper, we propose a Geometry Uncertainty Projection Network (GUP Net) to tackle the error amplification problem at both inference and training stages. Specifically, a GUP module is proposed to obtains the geometryguided uncertainty of the inferred depth, which not only provides high reliable confidence for each depth but also benefits depth learning. Furthermore, at the training stage, we propose a Hierarchical Task Learning strategy to reduce the instability caused by error amplification. This learning algorithm monitors the learning situation of each task by a proposed indicator and adaptively assigns the proper loss weights for different tasks according to their pre-tasks situation. Based on that, each task starts learning only when its pre-tasks are learned well, which can significantly improve the stability and efficiency of the training process. Extensive experiments demonstrate the effectiveness of the proposed method. The overall model can infer more reliable object depth than existing methods and outperforms the state-of-the-art image-based monocular 3D detectors by 3.74% and 4.7% AP 40 of the car and pedestrian categories on the KITTI benchmark. The code and model will be released at https://github.com/SuperMHP/GUPNet . โ This work was done when Yan Lu was an intern at SenseTime. * Equal contribution. Corresponding author. Supplementary material link ๐ โ 3๐ ๐๐๐๐กโ โ 2๐ Figure 1. The main pipeline of our Geometry Uncertainty Projection module. The projection process is modeled by the uncertainty theory in the probability framework. The inference depths can be represented as a distribution so that can provide both accurate values and scores.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9bf3cb5f-8eec-40fd-ab99-da21d8cb8ecbCited by top-tier papers71
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia et al.NeurIPS 2022 ยท 762 citations
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 ยท 199 citations
- Learning Auxiliary Monocular Contexts Helps Monocular 3D Object DetectionXianpeng Liu, Nan Xue, Tianfu WuAAAI 2022 ยท 181 citations
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo et al.ICCV 2023 ยท 175 citations
- MonoDistill: Learning Spatial Features for Monocular 3D Object DetectionZhiyu Chong, Xinzhu Ma, Hong Zhang, Yuxin Yue et al.ICLR 2022 ยท 125 citations
Builds on12
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 ยท 1,467 citations
- STD: Sparse-to-Dense 3D Object Detector for Point CloudZetong Yang, Yanan Sun, Shu Liu, Xiaoyong Shen et al.ICCV 2019 ยท 840 citations
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 ยท 542 citations
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulรฒ, Lorenzo Porzi, Manuel Lopez-Antequera et al.ICCV 2019 ยท 504 citations
- Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous DrivingXinzhu Ma, Zhihui Wang, Haojie Li, Pengbo Zhang et al.ICCV 2019 ยท 339 citations
Related papers
- MonoDGP: Monocular 3D Object Detection with Decoupled-Query and Geometry-Error PriorsFanqi Pu, Yifan Wang, Jiru Deng, Wenming YangCVPR 2025
- MonoJSG: Joint Semantic and Geometric Cost Volume for Monocular 3D Object DetectionQing Lian, Peiliang Li, Xiaozhi ChenCVPR 2022 ยท 69 citations
- MonoGround: Detecting Monocular 3D Objects from the GroundZequn Qin, Xi LiCVPR 2022 ยท 68 citations
- Out-of-Distribution Detection for Monocular Depth EstimationJulia Hornauer, Adrian Holzbock, Vasileios BelagiannisICCV 2023 ยท 7 citations
- MonoCD: Monocular 3D Object Detection with Complementary DepthsLongfei Yan, Pei Yan, Shengzhou Xiong, Xuanyu Xiang et al.CVPR 2024 ยท 52 citations
