Leveraging Imagery Data with Spatial Point Prior for Weakly Semi-supervised 3D Object Detection
Hongzhi Gao, Zheng Chen, Zehui Chen, Lin Chen, Jiaming Liu, Shanghang Zhang, Feng Zhao
Abstract
Training high-accuracy 3D detectors necessitates massive labeled 3D annotations with 7 degree-of-freedom, which is laborious and time-consuming. Therefore, the form of point annotations is proposed to offer significant prospects for practical applications in 3D detection, which is not only more accessible and less expensive but also provides strong spatial information for object localization. In this paper, we empirically discover that it is non-trivial to merely adapt Point-DETR to its 3D form, encountering two main bottlenecks: 1) it fails to encode strong 3D prior into the model, and 2) it generates low-quality pseudo labels in distant regions due to the extreme sparsity of LiDAR points. To overcome these challenges, we introduce Point-DETR3D, a teacher-student framework for weakly semi-supervised 3D detection, designed to fully capitalize on point-wise supervision within a constrained instance-wise annotation budget. Different from Point-DETR which encodes 3D positional information solely through a point encoder, we propose an explicit positional query initialization strategy to enhance the positional prior. Considering the low quality of pseudo labels at distant regions produced by the teacher model, we enhance the detector's perception by incorporating dense imagery data through a novel Cross-Modal Deformable RoI Fusion (D-RoI). Moreover, an innovative point-guided self-supervised learning technique is proposed to allow for fully exploiting point priors, even in student models. Extensive experiments on representative nuScenes dataset demonstrate our Point-DETR3D obtains significant improvements compared to previous works. Notably, with only 5% of labeled data, Point-DETR3D achieves over 90% performance of its fully supervised counterpart.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b581145-e8d4-4c08-93fc-726088b9bbb5Builds on21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia et al.NeurIPS 2022 · 762 citations
- DeepInteraction: 3D Object Detection via Modality InteractionZeyu Yang, Jiaqi Chen, Zhenwei Miao, Wei Li et al.NeurIPS 2022 · 268 citations
Related papers
- A Simple Vision Transformer for Weakly Semi-supervised 3D Object DetectionDingyuan Zhang, Dingkang Liang, Zhikang Zou, Jingyu Li et al.ICCV 2023 · 36 citations
- Eliminating Spatial Ambiguity for Weakly Supervised 3D Object Detection without Spatial LabelsHaizhuang Liu, Huimin Ma, Yilin Wang, Bochao Zou et al.ACM MM 2022 · 6 citations
- MixSup: Mixed-grained Supervision for Label-efficient LiDAR-based 3D Object DetectionYuxue Yang, Lue Fan, Zhaoxiang ZhangICLR 2024 · 11 citations
- WeakM3D: Towards Weakly Supervised Monocular 3D Object DetectionLiang Peng, Senbo Yan, Boxi Wu, Zheng Yang et al.ICLR 2022 · 25 citations
- Learning with Noisy Data for Semi-Supervised 3D Object DetectionZehui Chen, Zhenyu Li, Shuo Wang, Dengpan Fu et al.ICCV 2023 · 14 citations
