Leveraging Vision-Centric Multi-Modal Expertise for 3D Object Detection
Linyan Huang, Zhiqi Li, Chonghao Sima, Wenhai Wang, Jingdong Wang, Yu Qiao, Hongyang Li
Abstract
Current research is primarily dedicated to advancing the accuracy of camera-only 3D object detectors (apprentice) through the knowledge transferred from LiDARor multi-modal-based counterparts (expert). However, the presence of the domain gap between LiDAR and camera features, coupled with the inherent incompatibility in temporal fusion, significantly hinders the effectiveness of distillation-based enhancements for apprentices. Motivated by the success of uni-modal distillation, an apprentice-friendly expert model would predominantly rely on camera features, while still achieving comparable performance to multi-modal models. To this end, we introduce VCD, a framework to improve the camera-only apprentice model, including an apprentice-friendly multi-modal expert and temporal-fusion-friendly distillation supervision. The multi-modal expert VCD-E adopts an identical structure as that of the camera-only apprentice in order to alleviate the feature disparity, and leverages LiDAR input as a depth prior to reconstruct the 3D scene, achieving the performance on par with other heterogeneous multi-modal experts. Additionally, a fine-grained trajectory-based distillation module is introduced with the purpose of individually rectifying the motion misalignment for each object in the scene. With those improvements, our camera-only apprentice VCD-A sets new state-of-the-art on nuScenes with a score of 63.1% NDS. The code will be released at https://github.com/OpenDriveLab/Birds-eye-view-Perception .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 441b53dd-35eb-41dc-9f81-4575e9b12b4bCited by top-tier papers9
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu et al.ICLR 2024 · 395 citations
- VeXKD: The Versatile Integration of Cross-Modal Fusion and Knowledge Distillation for 3D PerceptionYuzhe Ji, Yijie Chen, Liuqing Yang, Rui Ding et al.NeurIPS 2024 · 11 citations
- RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal FusionGeonho Bang, Minjae Seong, Jisong Kim, Geunju Baek et al.ICCV 2025 · 6 citations
- MemDistill: Distilling LiDAR Knowledge into Memory for Camera-Only 3D Object DetectionDonghyeon Kwon, Youngseok Yoon, Hyeongseok Son, Suha KwakICCV 2025 · 1 citation
- Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object DetectionHaowen Zheng, Hu Zhu, Lu Deng, Weihao Gu et al.AAAI 2026
Builds on31
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with TransformersXuyang Bai, Zeyu Hu, Xinge Zhu, Qingqiu Huang et al.CVPR 2022 · 794 citations
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia et al.NeurIPS 2022 · 762 citations
Related papers
- UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird's-Eye ViewShengchao Zhou, Weizhou Liu, Chen Hu, Shuchang Zhou et al.CVPR 2023
- BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang et al.ICLR 2023 · 28 citations
- CRKD: Enhanced Camera-Radar Object Detection with Cross-Modality Knowledge DistillationLingjun Zhao, Jingyu Song, Katherine A. SkinnerCVPR 2024 · 21 citations
- SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object DetectionHaimei Zhao, Qiming Zhang, Shanshan Zhao, Zhe Chen et al.AAAI 2024 · 31 citations
- Boosting 3D Object Detection by Simulating Multimodality on Point CloudsWu Zheng, Mingxuan Hong, Li Jiang, Chi-Wing FuCVPR 2022 · 32 citations
