MonoDiff: Monocular 3D Object Detection and Pose Estimation with Diffusion Models
Yasiru Ranasinghe, Deepti Hegde, Vishal M. Patel
摘要
3D object detection and pose estimation from a single-view image is challenging due to the high uncertainty caused by the absence of 3D perception. As a solution, recent monocular 3D detection methods leverage additional modalities, such as stereo image pairs and LiDAR point clouds, to enhance image features at the expense of additional annotation costs. We propose using diffusion models to learn effective representations for monoc-ular 3D detection without additional modalities or training data. We present MonoDiff, a novel framework that em-ploys the reverse diffusion process to estimate 3D bounding box and orientation. But, considering the variability in bounding box sizes along different dimensions, it is inef-fective to sample noise from a standard Gaussian distribution. Hence, we adopt a Gaussian mixture model to sam-ple noise during the forward diffusion process and initialize the reverse diffusion process. Furthermore, since the diffusion model generates the 3D parameters for a given object image, we leverage 2D detection information to pro-vide additional supervision by maintaining the correspon-dence between 3D/2D projection. Finally, depending on the signal-to-noise ratio, we incorporate a dynamic weighting scheme to account for the level of uncertainty in the supervision by projection at different timesteps. MonoDiff outperforms current state-of-the-art monocular 3D detection methods on the KITTI and Waymo benchmarks without additional depth priors. MonoDiff project is available at: https://dylran.github.iolmonodiffgithub.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Girish Chandar Ganesan, Xiaoming LiuCVPR 2026 · 被引用 13 次
- Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent DiffusionWentao Qu, Guofeng Mei, Jing Wang, Yujiao Wu 等AAAI 2026 · 被引用 6 次
- Towards Intrinsic-Aware Monocular 3D Object DetectionZhihao Zhang, Abhinav Kumar, Xiaoming LiuCVPR 2026 · 被引用 5 次
- CHARM3R: Towards Unseen Camera Height Robust Monocular 3D DetectorAbhinav Kumar, Yuliang Guo, Zhihao Zhang, Xinyu Huang 等ICCV 2025 · 被引用 1 次
- Promptable 3-D Object Localization with Latent Diffusion ModelsCheng-Yao Hong, Li-Heng Wang, Tyng-Luh LiuNeurIPS 2025
它引用的顶会 Paper32
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang 等NeurIPS 2022 · 被引用 1,546 次
相关 Paper
- DiffPose: Toward More Reliable 3D Pose EstimationJia Gong, Lin Geng Foo, Zhipeng Fan, Qiuhong Ke 等CVPR 2023
- 6D-Diff: A Keypoint Diffusion Framework for 6D Object Pose EstimationLi Xu, Haoxuan Qu, Yujun Cai, Jun LiuCVPR 2024 · 被引用 29 次
- Weakly Supervised Monocular 3D Detection with a Single-View ImageXueying Jiang, Sheng Jin, Lewei Lu, Xiaoqin Zhang 等CVPR 2024
- Learning Auxiliary Monocular Contexts Helps Monocular 3D Object DetectionXianpeng Liu, Nan Xue, Tianfu WuAAAI 2022 · 被引用 181 次
- MonoDTR: Monocular 3D Object Detection with Depth-Aware TransformerKuan-Chih Huang, Tsung-Han Wu, Hung-Ting Su, Winston H. HsuCVPR 2022 · 被引用 199 次
