STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object Detection
Sujin Jang, Dae Ung Jo, Sung Ju Hwang, Dongwook Lee, Daehyun Ji
摘要
3D object detection (3DOD) from multi-view images is an economically appealing alternative to expensive LiDAR-based detectors, but also an extremely challenging task due to the absence of precise spatial cues. Recent studies have leveraged the teacher-student paradigm for cross-modal distillation, where a strong LiDAR-modality teacher transfers useful knowledge to a multi-view-based image-modality student. However, prior approaches have only focused on minimizing global distances between cross-modal features, which may lead to suboptimal knowledge distillation results. Based on these insights, we propose a novel structural and temporal cross-modal knowledge distillation (STXD) framework for multi-view 3DOD. First, STXD reduces redundancy of the feature components of the student by regularizing the cross-correlation of cross-modal features, while maximizing their similarities. Second, to effectively transfer temporal knowledge, STXD encodes temporal relations of features across a sequence of frames via similarity maps. Lastly, STXD also adopts a response distillation method to further enhance the quality of knowledge distillation at the output-level. Our extensive experiments demonstrate that STXD significantly improves the NDS and mAP of the based student detectors by 2 . 8% ∼ 4 . 5% on the nuScenes testing dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Unified Domain Generalization and Adaptation for Multi-View 3D Object DetectionGyusam Chang, Jiwon Lee, Donghyun Kim, Jinkyu Kim 等NeurIPS 2024 · 被引用 19 次
- VRM: Knowledge Distillation via Virtual Relation MatchingWeijia Zhang, Fei Xie, Tom Weidong Cai, Chao MaICCV 2025 · 被引用 6 次
- RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal FusionGeonho Bang, Minjae Seong, Jisong Kim, Geunju Baek 等ICCV 2025 · 被引用 6 次
- Depth Any Event Stream: Enhancing Event-based Monocular Depth Estimation via Dense-to-Sparse DistillationJinjing Zhu, Tianbo Pan, Zidong Cao, Yexin Liu 等ICCV 2025 · 被引用 3 次
- Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object DetectionHaowen Zheng, Hu Zhu, Lu Deng, Weihao Gu 等AAAI 2026
它引用的顶会 Paper36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 被引用 1,226 次
- Voxel R-CNN: Towards High Performance Voxel-based 3D Object DetectionJiajun Deng, Shaoshuai Shi, Peiwei Li, Wengang Zhou 等AAAI 2021 · 被引用 1,128 次
相关 Paper
- MemDistill: Distilling LiDAR Knowledge into Memory for Camera-Only 3D Object DetectionDonghyeon Kwon, Youngseok Yoon, Hyeongseok Son, Suha KwakICCV 2025 · 被引用 1 次
- BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object DetectionZehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang 等ICLR 2023 · 被引用 28 次
- X3KD: Knowledge Distillation Across Modalities, Tasks and Stages for Multi-Camera 3D Object DetectionMarvin Klingner, Shubhankar Borse, Varun Ravi Kumar, Behnaz Rezaei 等CVPR 2023
- Boosting 3D Object Detection by Simulating Multimodality on Point CloudsWu Zheng, Mingxuan Hong, Li Jiang, Chi-Wing FuCVPR 2022 · 被引用 32 次
- StereoDistill: Pick the Cream from LiDAR for Distilling Stereo-Based 3D Object DetectionZhe Liu, Xiaoqing Ye, Xiao Tan, Errui Ding 等AAAI 2023 · 被引用 15 次
