DiffCAD: Weakly-Supervised Probabilistic CAD Model Retrieval and Alignment from an RGB Image
Daoyi Gao, Dávid Rozenberszki, Stefan Leutenegger, Angela Dai
摘要
Perceiving 3D structures from RGB images based on CAD model primitives can enable an effective, efficient 3D object-based representation of scenes. However, current approaches rely on supervision from expensive yet imperfect annotations of CAD models associated with real images, and encounter challenges due to the inherent ambiguities in the task - both in depth-scale ambiguity in monocular perception, as well as inexact matches of CAD database models to real observations. We thus propose DiffCAD, the first weakly-supervised probabilistic approach to CAD retrieval and alignment from an RGB image. We learn a probabilistic model through diffusion, modeling likely distributions of shape, pose, and scale of CAD objects in an image. This enables multi-hypothesis generation of different plausible CAD reconstructions, requiring only a few hypotheses to characterize ambiguities in depth/scale and inexact shape matches. Our approach is trained only on synthetic data, leveraging monocular depth and mask estimates to enable robust zero-shot adaptation to various real target domains. Despite being trained solely on synthetic data, our multi-hypothesis approach can even surpass the supervised state-of-the-art on the Scan2CAD dataset by 5.9% with 8 hypotheses.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion TransformersYuchen Lin, Chenguo Lin, Panwang Pan, Honglei Yan 等NeurIPS 2025 · 被引用 89 次
- LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D ScansZhening Huang, Xiaoyang Wu, Fangcheng Zhong, Hengshuang Zhao 等NeurIPS 2025 · 被引用 26 次
- WorldGen: From Text to Traversable and Interactive 3D WorldsDilin Wang, Hyunyoung Jung, Tom Monnier, Kihyuk Sohn 等CVPR 2026 · 被引用 24 次
- SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation ModelYukai Shi, Weiyu Li, Zihao Wang, Hongyang Li 等CVPR 2026 · 被引用 17 次
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single ImageZe-Xin Yin, Liu Liu, Xinjie wang, Wei Sui 等CVPR 2026 · 被引用 10 次
它引用的顶会 Paper43
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
相关 Paper
- Patch2CAD: Patchwise Embedding Learning for In-the-Wild Shape Retrieval from a Single ImageWeicheng Kuo, Anelia Angelova, Tsung-Yi Lin, Angela DaiICCV 2021 · 被引用 42 次
- Learning Local RGB-to-CAD Correspondences for Object Pose EstimationGeorgios Georgakis, Srikrishna Karanam, Ziyan Wu, Jana KoseckaICCV 2019 · 被引用 25 次
- Zero-Shot Inexact CAD Model Alignment from a Single ImagePattaramanee Arsomngern, Sasikarn Khwanmuang, Matthias Nießner, Supasorn SuwajanakornICCV 2025 · 被引用 2 次
- Diorama: Unleashing Zero-Shot Single-View 3D Indoor Scene ModelingQirui Wu, Denys Iliash, Daniel Ritchie, Manolis Savva 等ICCV 2025 · 被引用 4 次
- CADDreamer: CAD Object Generation from Single-view ImagesYuan Li, Cheng Lin, Yuan Liu, Xiaoxiao Long 等CVPR 2025
