Dimension Embeddings for Monocular 3D Object Detection
Yunpeng Zhang, Wenzhao Zheng, Zheng Zhu, Guan Huang, Dalong Du, Jie Zhou, Jiwen Lu
Abstract
Most existing deep learning-based approaches for monocular 3D object detection directly regress the dimensions of objects and overlook their importance in solving the illposed problem. In this paper, we propose a general method to learn appropriate embeddings for dimension estimation in monocular 3D object detection. Specifically, we consider two intuitive clues in learning the dimension-aware embeddings with deep neural networks. First, we constrain the pair-wise distance on the embedding space to reflect the similarity of corresponding dimensions so that the model can take advantage of inter-object information to learn more discriminative embeddings for dimension estimation. Second, we propose to learn representative shape templates on the dimension-aware embedding space. Through the attention mechanism, each object can interact with the learnable templates and obtain the attentive dimensions as the initial estimation, which is further refined by the combined features from both the object and the attentive templates. Experimental results on the well-established KITTI dataset demonstrate the proposed method of dimension embeddings can bring consistent improvements with negligible computation cost overhead. We achieve new state-of-the-art performance on the KITTI 3D object detection benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ebbccc9-3875-433d-8914-c855be125c52Cited by top-tier papers7
- MonoCD: Monocular 3D Object Detection with Complementary DepthsLongfei Yan, Pei Yan, Shengzhou Xiong, Xuanyu Xiang et al.CVPR 2024 · 52 citations
- Monocular 3D Object Detection with Bounding Box Denoising in 3D by PerceiverXianpeng Liu, Ce Zheng, Kelvin Cheng, Nan Xue et al.ICCV 2023 · 12 citations
- Multi-View Attentive Contextualization for Multi-View 3D Object DetectionXianpeng Liu, Ce Zheng, Ming Qian, Nan Xue et al.CVPR 2024 · 5 citations
- Towards Fair and Comprehensive Comparisons for Image-Based 3D Object DetectionXinzhu Ma, Yongtao Wang, Yinmin Zhang, Zhiyi Xia et al.ICCV 2023 · 1 citation
- NeurOCS: Neural NOCS Supervision for Monocular 3D Object LocalizationZhixiang Min, Bingbing Zhuang, Samuel Schulter, Buyu Liu et al.CVPR 2023
Builds on20
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera et al.ICCV 2019 · 504 citations
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li et al.ICCV 2021 · 404 citations
- Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous DrivingXinzhu Ma, Zhihui Wang, Haojie Li, Pengbo Zhang et al.ICCV 2019 · 339 citations
- Geometry Uncertainty Projection Network for Monocular 3D Object DetectionYan Lu, Xinzhu Ma, Lei Yang, Tianzhu Zhang et al.ICCV 2021 · 294 citations
Related papers
- AutoShape: Real-Time Shape-Aware Monocular 3D Object DetectionZongdai Liu, Dingfu Zhou, Feixiang Lu, Jin Fang et al.ICCV 2021 · 176 citations
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo et al.ICCV 2023 · 175 citations
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 199 citations
- Attention-Based Depth Distillation with 3D-Aware Positional Encoding for Monocular 3D Object DetectionZizhang Wu, Yunzhe Wu, Jian Pu, Xianzhi Li et al.AAAI 2023 · 29 citations
- M3DSSD: Monocular 3D Single Stage Object DetectorShujie Luo, Hang Dai, Ling Shao, Yong DingCVPR 2021
