Weakly Supervised Monocular 3D Object Detection Using Multi-View Projection and Direction Consistency
Runzhou Tao, Wencheng Han, Zhongying Qiu, Cheng-Zhong Xu, Jianbing Shen
Abstract
Monocular 3D object detection has become a mainstream approach in automatic driving for its easy application. A prominent advantage is that it does not need Li-DAR point clouds during the inference. However, most current methods still rely on 3D point cloud data for labeling the ground truths used in the training phase. This inconsistency between the training and inference makes it hard to utilize the large-scale feedback data and increases the data collection expenses. To bridge this gap, we propose a new weakly supervised monocular 3D objection detection method, which can train the model with only 2D labels marked on images. To be specific, we explore three types of consistency in this task, i.e. the projection, multi-view and direction consistency, and design a weakly-supervised architecture based on these consistencies. Moreover, we propose a new 2D direction labeling method in this task to guide the model for accurate rotation direction prediction. Experiments show that our weakly-supervised method achieves comparable performance with some fully supervised methods. When used as a pre-training method, our model can significantly outperform the corresponding fullysupervised baseline with only 1/3 3D labels.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 14e7093e-050e-44b6-a2bf-556e19bff2daCited by top-tier papers5
- Training an Open-Vocabulary Monocular 3D Detection Model without 3D DataRui Huang, Henry Zheng, Yan Wang, Zhuofan Xia et al.NeurIPS 2024 · 26 citations
- OLiDM: Object-aware LiDAR Diffusion Models for Autonomous DrivingTianyi Yan, Junbo Yin, Xianpeng Lang, Ruigang Yang et al.AAAI 2025 · 16 citations
- MonoSOWA: Scalable Monocular 3D Object Detector Without Human AnnotationsJan Skvrna, Lukás NeumannICCV 2025 · 3 citations
- Weakly Supervised Monocular 3D Detection with a Single-View ImageXueying Jiang, Sheng Jin, Lewei Lu, Xiaoqin Zhang et al.CVPR 2024
- VSRD: Instance-Aware Volumetric Silhouette Rendering for Weakly Supervised 3D Object DetectionZihua Liu, Hiroki Sakuma, Masatoshi OkutomiCVPR 2024
Builds on22
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
- Pseudo-LiDAR++: Accurate Depth for 3D Object Detection in Autonomous DrivingYurong You, Yan Wang, Wei-Lun Chao, Divyansh Garg et al.ICLR 2020 · 439 citations
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li et al.ICCV 2021 · 404 citations
- Geometry Uncertainty Projection Network for Monocular 3D Object DetectionYan Lu, Xinzhu Ma, Lei Yang, Tianzhu Zhang et al.ICCV 2021 · 294 citations
Related papers
- WeakM3D: Towards Weakly Supervised Monocular 3D Object DetectionLiang Peng, Senbo Yan, Boxi Wu, Zheng Yang et al.ICLR 2022 · 25 citations
- Eliminating Spatial Ambiguity for Weakly Supervised 3D Object Detection without Spatial LabelsHaizhuang Liu, Huimin Ma, Yilin Wang, Bochao Zou et al.ACM MM 2022 · 6 citations
- Weakly-Supervised 3D Human Pose Learning via Multi-View Images in the WildUmar Iqbal, Pavlo Molchanov, Jan KautzCVPR 2020
- H2RBox: Horizontal Box Annotation is All You Need for Oriented Object DetectionXue Yang, Gefan Zhang, Wentong Li, Yue Zhou et al.ICLR 2023 · 24 citations
- IDA-3D: Instance-Depth-Aware 3D Object Detection From Stereo Vision for Autonomous DrivingWanli Peng, Hao Pan, He Liu, Yi SunCVPR 2020
