Learning Auxiliary Monocular Contexts Helps Monocular 3D Object Detection
Xianpeng Liu, Nan Xue, Tianfu Wu
摘要
Monocular 3D object detection aims to localize 3D bounding boxes in an input single 2D image. It is a highly challenging problem and remains open, especially when no extra information (e.g., depth, lidar and/or multi-frames) can be leveraged in training and/or inference. This paper proposes a simple yet effective formulation for monocular 3D object detection without exploiting any extra information. It presents the MonoCon method which learns Monocular Contexts, as auxiliary tasks in training, to help monocular 3D object detection. The key idea is that with the annotated 3D bounding boxes of objects in an image, there is a rich set of well-posed projected 2D supervision signals available in training, such as the projected corner keypoints and their associated offset vectors with respect to the center of 2D bounding box, which should be exploited as auxiliary tasks in training. The proposed MonoCon is motivated by the Cramer–Wold theorem in measure theory at a high level. In implementation, it utilizes a very simple end-to-end design to justify the effectiveness of learning auxiliary monocular contexts, which consists of three components: a Deep Neural Network (DNN) based feature backbone, a number of regression head branches for learning the essential parameters used in the 3D bounding box prediction, and a number of regression head branches for learning auxiliary contexts. After training, the auxiliary context regression branches are discarded for better inference efficiency. In experiments, the proposed MonoCon is tested in the KITTI benchmark (car, pedestrian and cyclist). It outperforms all prior arts in the leaderboard on the car category and obtains comparable performance on pedestrian and cyclist in terms of accuracy. Thanks to the simple design, the proposed MonoCon method obtains the fastest inference speed with 38.7 fps in comparisons. Our code is released at https://git.io/MonoCon.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- MonoUNI: A Unified Vehicle and Infrastructure-side Monocular 3D Object Detection Network with Sufficient Depth CluesJinrang Jia, Zhenjia Li, Yifeng ShiNeurIPS 2023 · 被引用 69 次
- MonoCD: Monocular 3D Object Detection with Complementary DepthsLongfei Yan, Pei Yan, Shengzhou Xiong, Xuanyu Xiang 等CVPR 2024 · 被引用 52 次
- MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked AutoencodersXueying Jiang, Sheng Jin, Xiaoqin Zhang, Ling Shao 等NeurIPS 2024 · 被引用 31 次
- Attention-Based Depth Distillation with 3D-Aware Positional Encoding for Monocular 3D Object DetectionZizhang Wu, Yunzhe Wu, Jian Pu, Xianzhi Li 等AAAI 2023 · 被引用 29 次
- Learning Occupancy for Monocular 3D Object DetectionLiang Peng, Junkai Xu, Haoran Cheng, Zheng Yang 等CVPR 2024 · 被引用 21 次
它引用的顶会 Paper20
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 被引用 1,467 次
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 被引用 542 次
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera 等ICCV 2019 · 被引用 504 次
- Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous DrivingXinzhu Ma, Zhihui Wang, Haojie Li, Pengbo Zhang 等ICCV 2019 · 被引用 339 次
- Geometry Uncertainty Projection Network for Monocular 3D Object DetectionYan Lu, Xinzhu Ma, Lei Yang, Tianzhu Zhang 等ICCV 2021 · 被引用 294 次
相关 Paper
- MonoDistill: Learning Spatial Features for Monocular 3D Object DetectionZhiyu Chong, Xinzhu Ma, Hong Zhang, Yuxin Yue 等ICLR 2022 · 被引用 125 次
- Monocular 3D Object Detection with Decoupled Structured Polygon Estimation and Height-Guided Depth EstimationYingjie Cai, Buyu Li, Zeyu Jiao, Hongsheng Li 等AAAI 2020 · 被引用 100 次
- MoNet3D: Towards Accurate Monocular 3D Object Localization in Real TimeXichuan Zhou, Yicong Peng, Chunqiao Long, Fengbo Ren 等ICML 2020 · 被引用 15 次
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 被引用 199 次
- MonoGround: Detecting Monocular 3D Objects from the GroundZequn Qin, Xi LiCVPR 2022 · 被引用 68 次
