MonoCD: Monocular 3D Object Detection with Complementary Depths
Longfei Yan, Pei Yan, Shengzhou Xiong, Xuanyu Xiang, Yihua Tan
Abstract
Monocular 3D object detection has attracted widespread attention due to its potential to accurately obtain object 3D localization from a single image at a low cost. Depth estimation is an essential but challenging subtask of monocular 3D object detection due to the ill-posedness of 2D to 3D mapping. Many methods explore multiple local depth clues such as object heights and keypoints and then formulate the object depth estimation as an ensemble of multiple depth predictions to mitigate the insufficiency of single-depth information. However, the errors of existing multiple depths tend to have the same sign, which hinders them from neutralizing each other and limits the overall accuracy of combined depth. To alleviate this problem, we propose to increase the complementarity of depths with two novel designs. First, we add a new depth prediction branch named complementary depth that utilizes global and efficient depth clues from the entire image rather than the local clues to reduce the similarity of depth predictions. Second, we propose to fully exploit the geometric relations between multiple depth clues to achieve complementarity in form. Benefiting from these designs, our method achieves higher complementarity. Experiments on the KITTI bench-mark demonstrate that our method achieves state-of-the-art performance without introducing extra data. In addition, complementary depth can also be a lightweight and plug-and-play module to boost multiple existing monocular 3d object detectors. Code is available at https://github.com/elvintanhust/MonoCD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8263abe0-664f-4e9f-b756-b36cfb48b54bCited by top-tier papers18
- MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked AutoencodersXueying Jiang, Sheng Jin, Xiaoqin Zhang, Ling Shao et al.NeurIPS 2024 · 31 citations
- Training an Open-Vocabulary Monocular 3D Detection Model without 3D DataRui Huang, Henry Zheng, Yan Wang, Zhuofan Xia et al.NeurIPS 2024 · 26 citations
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan, Xiaoming LiuCVPR 2026 · 13 citations
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 5 citations
- MonoSOWA: Scalable Monocular 3D Object Detector Without Human AnnotationsJan Skvrna, Lukás NeumannICCV 2025 · 3 citations
Builds on23
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera et al.ICCV 2019 · 504 citations
- Geometry Uncertainty Projection Network for Monocular 3D Object DetectionYan Lu, Xinzhu Ma, Lei Yang, Tianzhu Zhang et al.ICCV 2021 · 294 citations
- How Do Neural Networks See Depth in Single Images?Tom van Dijk, Guido de CroonICCV 2019 · 210 citations
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 199 citations
- Behind the Curtain: Learning Occluded Shapes for 3D Object DetectionQiangeng Xu, Yiqi Zhong, Ulrich NeumannAAAI 2022 · 188 citations
Related papers
- Diversity Matters: Fully Exploiting Depth Clues for Reliable Monocular 3D Object DetectionZhuoling Li, Zhan Qu, Yang Zhou, Jianzhuang Liu et al.CVPR 2022 · 74 citations
- Objects Are Different: Flexible Monocular 3D Object DetectionYunpeng Zhang, Jiwen Lu, Jie ZhouCVPR 2021
- Attention-Based Depth Distillation with 3D-Aware Positional Encoding for Monocular 3D Object DetectionZizhang Wu, Yunzhe Wu, Jian Pu, Xianzhi Li et al.AAAI 2023 · 29 citations
- MonoGround: Detecting Monocular 3D Objects from the GroundZequn Qin, Xi LiCVPR 2022 · 68 citations
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo et al.ICCV 2023 · 175 citations
