Attention-Based Depth Distillation with 3D-Aware Positional Encoding for Monocular 3D Object Detection
Zizhang Wu, Yunzhe Wu, Jian Pu, Xianzhi Li, Xiaoquan Wang
摘要
Monocular 3D object detection is a low-cost but challenging task, as it requires generating accurate 3D localization solely from a single image input. Recent developed depth-assisted methods show promising results by using explicit depth maps as intermediate features, which are either precomputed by monocular depth estimation networks or jointly evaluated with 3D object detection. However, inevitable errors from estimated depth priors may lead to misaligned semantic information and 3D localization, hence resulting in feature smearing and suboptimal predictions. To mitigate this issue, we propose ADD, an Attention-based Depth knowledge Distillation framework with 3D-aware positional encoding. Unlike previous knowledge distillation frameworks that adopt stereo- or LiDAR-based teachers, we build up our teacher with identical architecture as the student but with extra ground-truth depth as input. Credit to our teacher design, our framework is seamless, domain-gap free, easily implementable, and is compatible with object-wise ground-truth depth. Specifically, we leverage intermediate features and responses for knowledge distillation. Considering long-range 3D dependencies, we propose 3D-aware self-attention and target-aware cross-attention modules for student adaptation. Extensive experiments are performed to verify the effectiveness of our framework on the challenging KITTI 3D object detection benchmark. We implement our framework on three representative monocular detectors, and we achieve state-of-the-art performance with no additional inference computational cost relative to baseline models. Our code is available at https://github.com/rockywind/ADD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- MonoMAE: Enhancing Monocular 3D Detection through Depth-Aware Masked AutoencodersXueying Jiang, Sheng Jin, Xiaoqin Zhang, Ling Shao 等NeurIPS 2024 · 被引用 31 次
- Depth-discriminative Metric Learning for Monocular 3D Object DetectionWonhyeok Choi, Mingyu Shin, Sunghoon ImNeurIPS 2023 · 被引用 13 次
- Monocular Vehicle Pose and Shape Reconstruction via Dynamic Context Adaptation and Progressive Geometry RefinementWei Li, Long Ji, Ying Wang, Xiao Wu 等AAAI 2026
- RARE: Learn to RAnk and REtrieve for Monocular 3D Object DetectionHyeonjeong Park, Peixi Xiong, Xiaoqian Ruan, Dian Jia 等CVPR 2026
- MonoTAKD: Teaching Assistant Knowledge Distillation for Monocular 3D Object DetectionHou-I Liu, Christine Wu, Jen-Hao Cheng, Wenhao Chai 等CVPR 2025
它引用的顶会 Paper25
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
- M3D-RPN: Monocular 3D Region Proposal Network for Object DetectionGarrick Brazil, Xiaoming LiuICCV 2019 · 被引用 542 次
- PETRv2: A Unified Framework for 3D Perception from Multi-Camera ImagesYingfei Liu, Junjie Yan, Fan Jia, Shuailin Li 等ICCV 2023 · 被引用 513 次
- Is Pseudo-Lidar needed for Monocular 3D Object detection?Dennis Park, Rares Ambrus, Vitor Guizilini, Jie Li 等ICCV 2021 · 被引用 404 次
相关 Paper
- Beyond the limitation of monocular 3D detector via knowledge distillationYiran Yang, Dongshuo Yin, Xuee Rong, Xian Sun 等ICCV 2023 · 被引用 4 次
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 被引用 199 次
- Weakly Supervised Monocular 3D Detection with a Single-View ImageXueying Jiang, Sheng Jin, Lewei Lu, Xiaoqin Zhang 等CVPR 2024
- FD3D: Exploiting Foreground Depth Map for Feature-Supervised Monocular 3D Object DetectionZizhang Wu, Yuanzhu Gan, Yunzhe Wu, Ruihao Wang 等AAAI 2024 · 被引用 19 次
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo 等ICCV 2023 · 被引用 175 次
