SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection
Yifan Wang, Yian Zhao, Fanqi Pu, Xiaochen Yang, Yang Tang, Xi Chen, Wenming Yang
摘要
Existing monocular 3D detectors typically tame the pronounced nonlinear regression of 3D bounding box through decoupled prediction paradigm, which employs multiple branches to estimate geometric center, depth, dimensions, and rotation angle separately. Although this decoupling strategy simplifies the learning process, it inherently ignores the geometric collaborative constraints between different attributes, resulting in the lack of geometric consistency prior, thereby leading to suboptimal performance. To address this issue, we propose novel Spatial-Projection Alignment (SPAN) with two pivotal components: (i). Spatial Point Alignment enforces an explicit global spatial constraint between the predicted and ground-truth 3D bounding boxes, thereby rectifying spatial drift caused by decoupled attribute regression. (ii). 3D-2D Projection Alignment ensures that the projected 3D box is aligned tightly within its corresponding 2D detection bounding box on the image plane, mitigating projection misalignment overlooked in previous works. To ensure training stability, we further introduce a Hierarchical Task Learning strategy that progressively incorporates spatial-projection alignment as 3D attribute predictions refine, preventing early stage error propagation across attributes. Extensive experiments demonstrate that the proposed method can be easily integrated into any established monocular 3D detector and delivers significant performance improvements. Project
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei 等CVPR 2024 · 被引用 3,046 次
- Disentangling Monocular 3D Object DetectionAndrea Simonelli, Samuel Rota Bulò, Lorenzo Porzi, Manuel Lopez-Antequera 等ICCV 2019 · 被引用 504 次
- Geometry Uncertainty Projection Network for Monocular 3D Object DetectionYan Lu, Xinzhu Ma, Lei Yang, Tianzhu Zhang 等ICCV 2021 · 被引用 294 次
- Learning Auxiliary Monocular Contexts Helps Monocular 3D Object DetectionXianpeng Liu, Nan Xue, Tianfu WuAAAI 2022 · 被引用 181 次
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo 等ICCV 2023 · 被引用 175 次
相关 Paper
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan, Xiaoming LiuCVPR 2026 · 被引用 13 次
- Monocular 3D Object Detection with Decoupled Structured Polygon Estimation and Height-Guided Depth EstimationYingjie Cai, Buyu Li, Zeyu Jiao, Hongsheng Li 等AAAI 2020 · 被引用 100 次
- Delving Into Localization Errors for Monocular 3D Object DetectionXinzhu Ma, Yinmin Zhang, Dan Xu, Dongzhan Zhou 等CVPR 2021
- AutoShape: Real-Time Shape-Aware Monocular 3D Object DetectionZongdai Liu, Dingfu Zhou, Feixiang Lu, Jin Fang 等ICCV 2021 · 被引用 176 次
- Geometry-based Distance Decomposition for Monocular 3D Object DetectionXuepeng Shi, Qi Ye, Xiaozhi Chen, Chuangrong Chen 等ICCV 2021 · 被引用 169 次
