PromptDet: A Lightweight 3D Object Detection Framework with LiDAR Prompts
Kun Guo, Qiang Ling
Abstract
Multi-camera 3D object detection aims to detect and localize objects in 3D space using multiple cameras, which has attracted more attention due to its cost-effectiveness trade-off. However, these methods often struggle with the lack of accurate depth estimation caused by the natural weakness of the camera in ranging. Recently, multi-modal fusion and knowledge distillation methods for 3D object detection have been proposed to solve this problem, which are time-consuming during the training phase and not friendly to memory cost. In light of this, we propose PromptDet, a lightweight yet effective 3D object detection framework motivated by the success of prompt learning in 2D foundation model. Our proposed framework, PromptDet, comprises two integral components: a general camera-based detection module, exemplified by models like BEVDet and BEVDepth, and a LiDAR-assisted prompter. The LiDAR-assisted prompter leverages the LiDAR points as a complementary signal, enriched with a minimal set of additional trainable parameters. Notably, our framework is flexible due to our prompt-like design, which can not only be used as a lightweight multi-modal fusion method but also as a camera-only method for 3D object detection during the inference phase. Extensive experiments on nuScenes validate the effectiveness of the proposed PromptDet. As a multi-modal detector, PromptDet improves the mAP and NDS by at most 22.8% and 21.1% with fewer than 2% extra parameters compared with the camera-only baseline. Without LiDAR points, PromptDet still achieves an improvement of at most 2.4% mAP and 4.0% NDS with almost no impact on camera detection inference time. We will release our code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on13
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia et al.NeurIPS 2022 · 762 citations
- FB-BEV: BEV Representation from Forward-Backward View TransformationsZhiqi Li, Zhiding Yu, Wenhai Wang, Anima Anandkumar et al.ICCV 2023 · 144 citations
- Cross Modal Transformer: Towards Fast and Robust 3D Object DetectionJunjie Yan, Yingfei Liu, Jianjian Sun, Fan Jia et al.ICCV 2023 · 143 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
Related papers
- BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang et al.ICLR 2023 · 28 citations
- MemDistill: Distilling LiDAR Knowledge into Memory for Camera-Only 3D Object DetectionDonghyeon Kwon, Youngseok Yoon, Hyeongseok Son, Suha KwakICCV 2025 · 1 citation
- SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object DetectionHaimei Zhao, Qiming Zhang, Shanshan Zhao, Zhe Chen et al.AAAI 2024 · 31 citations
- GAFusion: Adaptive Fusing LiDAR and Camera with Multiple Guidance for 3D Object DetectionXiaotian Li, Baojie Fan, Jiandong Tian, Huijie FanCVPR 2024
- Boosting 3D Object Detection by Simulating Multimodality on Point CloudsWu Zheng, Mingxuan Hong, Li Jiang, Chi-Wing FuCVPR 2022 · 32 citations
