Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
Zhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan, Xiaoming Liu
摘要
Monocular 3D detection (Mono3D) aims to infer 3D bounding boxes from a single RGB image. Without auxiliary sensors such as LiDAR, this task is inherently ill-posed since the 3D-to-2D projection introduces depth ambiguity. Previous works often predict 3D attributes (e.g., depth, size, and orientation) in parallel, overlooking that these attributes are inherently correlated through the 3D-to-2D projection. However, simply enforcing such correlations through sequential prediction can propagate errors across attributes, especially when objects are occluded or truncated, where inaccurate size or orientation predictions can further amplify depth errors. Therefore, neither parallel nor sequential prediction is optimal. In this paper, we propose Mono-CoP, an adaptive framework that learns when and how to leverage inter-attribute correlations with two complementary designs. A Chain-of-Prediction (CoP) explores interattribute correlations through feature-level learning, propagation, and aggregation, while an Uncertainty-Guided Selector (GS) dynamically switches between CoP and par-allel paradigms for each object based on the predicted uncertainty. By combining their strengths, MonoCoP achieves state-of-the-art performance on KITTI, nuScenes, and Waymo, significantly improving depth accuracy, particularly for distant objects. Code and models are publicly available at https://github.com/alanzhangcs/MonoCoP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu 等ICLR 2026 · 被引用 15 次
- FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human RecognitionJie Zhu, Xiao Guo, Yiyang Su, Anil K. Jain 等CVPR 2026 · 被引用 7 次
- EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot PersonalizationHaolan Xu, Keli Cheng, Lei Wang, Ning Bi 等CVPR 2026 · 被引用 5 次
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 被引用 5 次
- MonoVLM: Monocular 3D Visual Grounding with Vision Language ModelsHuaizhi Qu, Hossein Nourkhiz Mahjoub, Vaishnav Tadiparthi, Kwonjoon Lee 等CVPR 2026
它引用的顶会 Paper58
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei 等CVPR 2024 · 被引用 3,046 次
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
相关 Paper
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo 等ICCV 2023 · 被引用 175 次
- SPAN: Spatial-Projection Alignment for Monocular 3D Object DetectionYifan Wang, Yian Zhao, Fanqi Pu, Xiaochen Yang 等CVPR 2026
- Monocular 3D Object Detection with Decoupled Structured Polygon Estimation and Height-Guided Depth EstimationYingjie Cai, Buyu Li, Zeyu Jiao, Hongsheng Li 等AAAI 2020 · 被引用 100 次
- Learning Auxiliary Monocular Contexts Helps Monocular 3D Object DetectionXianpeng Liu, Nan Xue, Tianfu WuAAAI 2022 · 被引用 181 次
- MonoDGP: Monocular 3D Object Detection with Decoupled-Query and Geometry-Error PriorsFanqi Pu, Yifan Wang, Jiru Deng, Wenming YangCVPR 2025
