DistillBEV: Boosting Multi-Camera 3D Object Detection with Cross-Modal Knowledge Distillation
Zeyu Wang, Dingwen Li, Chenxu Luo, Cihang Xie, Xiaodong Yang
Abstract
3D perception based on the representations learned from multi-camera bird’s-eye-view (BEV) is trending as cameras are cost-effective for mass production in autonomous driving industry. However, there exists a distinct performance gap between multi-camera BEV and LiDAR based 3D object detection. One key reason is that LiDAR captures accurate depth and other geometry measurements, while it is notoriously challenging to infer such 3D information from merely image input. In this work, we propose to boost the representation learning of a multi-camera BEV based student detector by training it to imitate the features of a well-trained LiDAR based teacher detector. We propose effective balancing strategy to enforce the student to focus on learning the crucial features from the teacher, and generalize knowledge transfer to multi-scale layers with temporal fusion. We conduct extensive evaluations on multiple representative models of multi-camera BEV. Experiments reveal that our approach renders significant improvement over the student models, leading to the state-of-the-art performance on the popular benchmark nuScenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- SCKD: Semi-Supervised Cross-Modality Knowledge Distillation for 4D Radar Object DetectionRuoyu Xu, Zhiyu Xiang, Chenwei Zhang, Hanzhi Zhong et al.AAAI 2025 · 27 citations
- CRKD: Enhanced Camera-Radar Object Detection with Cross-Modality Knowledge DistillationLingjun Zhao, Jingyu Song, Katherine A. SkinnerCVPR 2024 · 21 citations
- DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous DrivingChen Min, Dawei Zhao, Liang Xiao, Jian Zhao et al.CVPR 2024 · 20 citations
- 3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene UnderstandingXiaohu Huang, Jingjing Wu, Qunyi Xie, Kai HanNeurIPS 2025 · 11 citations
- VeXKD: The Versatile Integration of Cross-Modal Fusion and Knowledge Distillation for 3D PerceptionYuzhe Ji, Yijie Chen, Liuqing Yang, Rui Ding et al.NeurIPS 2024 · 11 citations
Builds on20
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li et al.ICCV 2021 · 404 citations
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li et al.NeurIPS 2022 · 401 citations
- Multimodal Virtual Point 3D DetectionTianwei Yin, Xingyi Zhou, Philipp KrähenbühlNeurIPS 2021 · 379 citations
Related papers
- BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang et al.ICLR 2023 · 28 citations
- UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird's-Eye ViewShengchao Zhou, Weizhou Liu, Chen Hu, Shuchang Zhou et al.CVPR 2023
- GeoMIM: Towards Better 3D Knowledge Transfer via Masked Image Modeling for Multi-view 3D UnderstandingJihao Liu, Tai Wang, Boxiao Liu, Qihang Zhang et al.ICCV 2023 · 22 citations
- BEVDilation: LiDAR-Centric Multi-Modal Fusion for 3D Object DetectionGuowen Zhang, Chenhang He, Liyi Chen, Lei ZhangAAAI 2026 · 2 citations
- BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving ScenariosZhiwei Lin, Yongtao Wang, Shengxiang Qi, Nan Dong et al.AAAI 2024 · 32 citations
