SP3D: Boosting Sparsely-Supervised 3D Object Detection via Accurate Cross-Modal Semantic Prompts
Shijia Zhao, Qiming Xia, Xusheng Guo, Pufan Zou, Maoji Zheng, Hai Wu, Chenglu Wen, Cheng Wang
Abstract
Recently, sparsely-supervised 3D object detection has gained great attention, achieving performance close to fully-supervised 3D detectors while requiring only a few annotated instances. Nevertheless, these methods suffer challenges when accurate labels are extremely absent. In this paper, we propose a boosting strategy, termed SP3D, explicitly utilizing the cross-modal semantic prompts generated from Large Multimodal Models (LMMs) to boost the 3D detector with robust feature discrimination capability under sparse annotation settings. Specifically, we first develop a Confident Points Semantic Transfer (CPST) module that generates accurate cross-modal semantic prompts through boundary-constrained center cluster selection. Based on these accurate semantic prompts, which we treat as seed points, we introduce a Dynamic Cluster Pseudo-label Generation (DCPG) module to yield pseudo-supervision signals from the geometry shape of multi-scale neighbor points. Additionally, we design a Distribution Shape score (DS score) that chooses highquality supervision signals for the initial training of the 3D detector. Experiments on the KITTI dataset and Waymo Open Dataset (WOD) have validated that SP3D can enhance the performance of sparsely supervised detectors by a large margin under meager labeling conditions. Moreover, we verified SP3D in the zero-shot setting, where its performance exceeded that of the state-of-the-art methods. The code is available at https://github.com/ xmuqimingxia/SP3D.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68b8c653-b6c6-4be6-bac2-e31512768c1cCited by top-tier papers2
- MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated LabelJunyoung Jung, Seokwon Kim, Jung Uk KimCVPR 2026 · 1 citation
- TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object DetectionLeyuan Xing, huanjia zhang, Dongyu Pan, Hai Wu et al.CVPR 2026
Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou et al.AAAI 2021 · 1,128 citations
- PointCLIP V2: Prompting CLIP and GPT for Powerful 3D Open-world LearningXiangyang Zhu, Renrui Zhang, Bowei He, Ziyu Guo et al.ICCV 2023 · 248 citations
Related papers
- SS3D: Sparsely-Supervised 3D Object Detection from Point CloudChuandong Liu, Chenqiang Gao, Fangcen Liu, Jiang Liu et al.CVPR 2022 · 32 citations
- MWSIS: Multimodal Weakly Supervised Instance Segmentation with 2D Box Annotations for Autonomous DrivingGuangfeng Jiang, Jun Liu, Yuzhi Wu, Wenlong Liao et al.AAAI 2024 · 11 citations
- Weakly Supervised 3D Object Detection from Point CloudsZengyi Qin, Jinglu Wang, Yan LuACM MM 2020 · 68 citations
- Commonsense Prototype for Outdoor Unsupervised 3D Object DetectionHai Wu, Shijia Zhao, Xun Huang, Chenglu Wen et al.CVPR 2024 · 15 citations
- Leveraging Imagery Data with Spatial Point Prior for Weakly Semi-supervised 3D Object DetectionHongzhi Gao, Zheng Chen, Zehui Chen, Lin Chen et al.AAAI 2024 · 3 citations
