UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird's-Eye View
Shengchao Zhou, Weizhou Liu, Chen Hu, Shuchang Zhou, Chao Ma
Abstract
In the field of 3D object detection for autonomous driving, the sensor portfolio including multi-modality and single-modality is diverse and complex. Since the multimodal methods have system complexity while the accuracy of single-modal ones is relatively low, how to make a tradeoff between them is difficult. In this work, we propose a universal cross-modality knowledge distillation framework (UniDistill) to improve the performance of single-modality detectors. Specifically, during training, UniDistill projects the features of both the teacher and the student detector into Bird's-Eye-View (BEV), which is a friendly representation for different modalities. Then, three distillation losses are calculated to sparsely align the foreground features, helping the student learn from the teacher without introducing additional cost during inference. Taking advantage of the similar detection paradigm of different detectors in BEV, UniDistill easily supports LiDAR-to-camera, camera-to-LiDAR, fusion-to-LiDAR and fusion-to-camera distillation paths. Furthermore, the three distillation losses can filter the effect of misaligned background information and balance between objects of different sizes, improving the distillation effectiveness. Extensive experiments on nuScenes demonstrate that UniDistill effectively improves the mAP and NDS of student detectors by 2.0%∼3.2%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- SparseOcc: Rethinking Sparse Latent Representation for Vision-Based Semantic Occupancy PredictionPin Tang, Zhongdao Wang, Guoqing Wang, Jilai Zheng et al.CVPR 2024 · 37 citations
- SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object DetectionHaimei Zhao, Qiming Zhang, Shanshan Zhao, Zhe Chen et al.AAAI 2024 · 31 citations
- SCKD: Semi-Supervised Cross-Modality Knowledge Distillation for 4D Radar Object DetectionRuoyu Xu, Zhiyu Xiang, Chenwei Zhang, Hanzhi Zhong et al.AAAI 2025 · 27 citations
- Not All Voxels are Equal: Hardness-Aware Semantic Scene Completion with Self-DistillationSong Wang, Jiawei Yu, Wentong Li, Wenyu Liu et al.CVPR 2024 · 22 citations
- CRKD: Enhanced Camera-Radar Object Detection with Cross-Modality Knowledge DistillationLingjun Zhao, Jingyu Song, Katherine A. SkinnerCVPR 2024 · 21 citations
Builds on25
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 542 citations
- Voxel Transformer for 3D Object DetectionJiageng Mao, Yujing Xue, Minzhe Niu, Haoyue Bai et al.ICCV 2021 · 535 citations
- TANet: Robust 3D Object Detection from Point Clouds with Triple AttentionZhe Liu, Xin Zhao, Tengteng Huang, Ruolan Hu et al.AAAI 2020 · 412 citations
Related papers
- BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang et al.ICLR 2023 · 28 citations
- MemDistill: Distilling LiDAR Knowledge into Memory for Camera-Only 3D Object DetectionDonghyeon Kwon, Youngseok Yoon, Hyeongseok Son, Suha KwakICCV 2025 · 1 citation
- Boosting 3D Object Detection by Simulating Multimodality on Point CloudsWu Zheng, Mingxuan Hong, Li Jiang, Chi-Wing FuCVPR 2022 · 32 citations
- X3KD: Knowledge Distillation Across Modalities, Tasks and Stages for Multi-Camera 3D Object DetectionMarvin Klingner, Shubhankar Borse, Varun Ravi Kumar, Behnaz Rezaei et al.CVPR 2023
- STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object DetectionSujin Jang, Dae Ung Jo, Sung Ju Hwang, Dongwook Lee et al.NeurIPS 2023 · 19 citations
