Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object Detection
Haowen Zheng, Hu Zhu, Lu Deng, Weihao Gu, Yang Yang, Yanyan Liang
Abstract
Camera-based temporal 3D object detection has shown impressive results in autonomous driving, with offline models improving accuracy by using future frames. Knowledge distillation (KD) can be an appealing framework for transferring rich information from offline models to online models. However, existing KD methods overlook future frames, as they mainly focus on spatial feature distillation under strict frame alignment or on temporal relational distillation, thereby making it challenging for online models to effectively learn future knowledge. To this end, we propose a sparse query-based approach, Future Temporal Knowledge Distillation (FTKD), which effectively transfers future frame knowledge from an offline teacher model to an online student model. Specifically, we present a future-aware feature reconstruction strategy to encourage the student model to capture future features without strict frame alignment. In addition, we further introduce future-guided logit distillation to leverage the teacher's stable foreground and background context. FTKD is applied to two high-performing 3D object detection baselines, achieving up to 1.3 mAP and 1.3 NDS gains on the nuScenes dataset, as well as the most accurate velocity estimation, without increasing inference cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 015d1f03-eaf1-4f92-9284-9b567f8efc05Builds on20
- BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object DetectionYinhao Li, Zheng Ge, Guanyi Yu, Jinrong Yang et al.AAAI 2023 · 954 citations
- Channel-wise Knowledge Distillation for Dense Prediction*Changyong Shu, Yifan Liu, Jianfei Gao, Zheng Yan et al.ICCV 2021 · 432 citations
- Unifying Voxel-based Representation with Transformer for 3D Object DetectionYanwei Li, Yilun Chen, Xiaojuan Qi, Zeming Li et al.NeurIPS 2022 · 401 citations
- Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object DetectionShihao Wang, Yingfei Liu, Tiancai Wang, Ying Li et al.ICCV 2023 · 399 citations
- Improve Object Detection with Feature-based Knowledge Distillation: Towards Accurate and Efficient DetectorsLinfeng Zhang, Kaisheng MaICLR 2021 · 251 citations
Related papers
- STXD: Structural and Temporal Cross-Modal Distillation for Multi-View 3D Object DetectionSujin Jang, Dae Ung Jo, Sung Ju Hwang, Dongwook Lee et al.NeurIPS 2023 · 19 citations
- RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal FusionGeonho Bang, Minjae Seong, Jisong Kim, Geunju Baek et al.ICCV 2025 · 6 citations
- MemDistill: Distilling LiDAR Knowledge into Memory for Camera-Only 3D Object DetectionDonghyeon Kwon, Youngseok Yoon, Hyeongseok Son, Suha KwakICCV 2025 · 1 citation
- MonoTAKD: Teaching Assistant Knowledge Distillation for Monocular 3D Object DetectionHou-I Liu, Christine Wu, Jen-Hao Cheng, Wenhao Chai et al.CVPR 2025
- Query-based Temporal Fusion with Explicit Motion for 3D Object DetectionJinghua Hou, Zhe Liu, Dingkang Liang, Zhikang Zou et al.NeurIPS 2023 · 28 citations
