Patch2CAD: Patchwise Embedding Learning for In-the-Wild Shape Retrieval from a Single Image
Weicheng Kuo, Anelia Angelova, Tsung-Yi Lin, Angela Dai
Abstract
3D perception of object shapes from RGB image input is fundamental towards semantic scene understanding, grounding image-based perception in our spatially 3dimensional real-world environments. To achieve a mapping between image views of objects and 3D shapes, we leverage CAD model priors from existing large-scale databases, and propose a novel approach towards constructing a joint embedding space between 2D images and 3D CAD models in a patch-wise fashion – establishing correspondences between patches of an image view of an object and patches of CAD geometry. This enables part similarity reasoning for retrieving similar CADs to a new image view without exact matches in the database. Our patch embedding provides more robust CAD retrieval for shape estimation in our end-to-end estimation of CAD model shape and pose for detected objects in a single input image. Experiments on in-the-wild, complex imagery from ScanNet show that our approach is more robust than state of the art in real-world scenarios without any exact CAD matches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 33b03779-7427-4bf3-9cfe-d241affbdfe1Cited by top-tier papers18
- ROCA: Robust CAD Model Retrieval and Alignment from a Single ImageCan Gümeli, Angela Dai, Matthias NießnerCVPR 2022 · 43 citations
- Mesh2Tex: Generating Mesh Textures from Image QueriesAlexey Bokhovkin, Shubham Tulsiani, Angela DaiICCV 2023 · 32 citations
- CAST: Component-Aligned 3D Scene Reconstruction from an RGB ImageKaixin Yao, Longwen Zhang, Xinhao Yan, Yan Zeng et al.SIGGRAPH 2025 · 30 citations
- DiffCAD: Weakly-Supervised Probabilistic CAD Model Retrieval and Alignment from an RGB ImageDaoyi Gao, Dávid Rozenberszki, Stefan Leutenegger, Angela DaiSIGGRAPH 2024 · 28 citations
- WorldGen: From Text to Traversable and Interactive 3D WorldsDilin Wang, Hyunyoung Jung, Tom Monnier, Kihyuk Sohn et al.CVPR 2026 · 24 citations
Builds on11
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- PointFlow: 3D Point Cloud Generation With Continuous Normalizing FlowsGuandao Yang, Xun Huang, Zekun Hao, Ming-Yu Liu et al.ICCV 2019 · 794 citations
- Pixel2Mesh++: Multi-View 3D Mesh Generation via DeformationChao Wen, Yinda Zhang, Zhuwen Li, Yanwei FuICCV 2019 · 279 citations
- End-to-End CAD Model Retrieval and 9DoF Alignment in 3D ScansArmen Avetisyan, Angela Dai, Matthias NießnerICCV 2019 · 88 citations
Related papers
- Learning Local RGB-to-CAD Correspondences for Object Pose EstimationGeorgios Georgakis, Srikrishna Karanam, Ziyan Wu, Jana KoseckaICCV 2019 · 25 citations
- Joint Embedding of 3D Scan and CAD ObjectsManuel Dahnert, Angela Dai, Leonidas J. Guibas, Matthias NießnerICCV 2019 · 36 citations
- Learning 3D Scene Priors with 2D SupervisionYinyu Nie, Angela Dai, Xiaoguang Han, Matthias NießnerCVPR 2023
- Zero-Shot Inexact CAD Model Alignment from a Single ImagePattaramanee Arsomngern, Sasikarn Khwanmuang, Matthias Nießner, Supasorn SuwajanakornICCV 2025 · 2 citations
- KP-RED: Exploiting Semantic Keypoints for Joint 3D Shape Retrieval and DeformationRuida Zhang, Chenyangguang Zhang, Yan Di, Fabian Manhardt et al.CVPR 2024 · 2 citations
