Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
Zhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan, Xiaoming Liu
Abstract
Monocular 3D detection (Mono3D) aims to infer 3D bounding boxes from a single RGB image. Without auxiliary sensors such as LiDAR, this task is inherently ill-posed since the 3D-to-2D projection introduces depth ambiguity. Previous works often predict 3D attributes (e.g., depth, size, and orientation) in parallel, overlooking that these attributes are inherently correlated through the 3D-to-2D projection. However, simply enforcing such correlations through sequential prediction can propagate errors across attributes, especially when objects are occluded or truncated, where inaccurate size or orientation predictions can further amplify depth errors. Therefore, neither parallel nor sequential prediction is optimal. In this paper, we propose Mono-CoP, an adaptive framework that learns when and how to leverage inter-attribute correlations with two complementary designs. A Chain-of-Prediction (CoP) explores interattribute correlations through feature-level learning, propagation, and aggregation, while an Uncertainty-Guided Selector (GS) dynamically switches between CoP and par-allel paradigms for each object based on the predicted uncertainty. By combining their strengths, MonoCoP achieves state-of-the-art performance on KITTI, nuScenes, and Waymo, significantly improving depth accuracy, particularly for distant objects. Code and models are publicly available at https://github.com/alanzhangcs/MonoCoP.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10e85c4c-78c5-4b00-b718-d86dba6ec09cCited by top-tier papers6
- Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model OutputsYiwei Chen, Soumyadeep Pal, Yimeng Zhang, Qing Qu et al.ICLR 2026 · 15 citations
- FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human RecognitionJie Zhu, Xiao Guo, Yiyang Su, Anil K. Jain et al.CVPR 2026 · 7 citations
- EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot PersonalizationHaolan Xu, Keli Cheng, Lei Wang, Ning Bi et al.CVPR 2026 · 5 citations
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 5 citations
- MonoVLM: Monocular 3D Visual Grounding with Vision Language ModelsHuaizhi Qu, Hossein Nourkhiz Mahjoub, Vaishnav Tadiparthi, Kwonjoon Lee et al.CVPR 2026
Builds on58
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
Related papers
- MonoDETR: Depth-guided Transformer for Monocular 3D Object DetectionRenrui Zhang, Han Qiu, Tai Wang, Ziyu Guo et al.ICCV 2023 · 175 citations
- SPAN: Spatial-Projection Alignment for Monocular 3D Object DetectionYifan Wang, Yian Zhao, Fanqi Pu, Xiaochen Yang et al.CVPR 2026
- Monocular 3D Object Detection with Decoupled Structured Polygon Estimation and Height-Guided Depth EstimationYingjie Cai, Buyu Li, Zeyu Jiao, Hongsheng Li et al.AAAI 2020 · 100 citations
- Learning Auxiliary Monocular Contexts Helps Monocular 3D Object DetectionXianpeng Liu, Nan Xue, Tianfu WuAAAI 2022 · 181 citations
- MonoDGP: Monocular 3D Object Detection with Decoupled-Query and Geometry-Error PriorsFanqi Pu, Yifan Wang, Jiru Deng, Wenming YangCVPR 2025
