CAD-Estate: Large-scale CAD Model Annotation in RGB Videos
Kevis-Kokitsi Maninis, Stefan Popov, Matthias Nießner, Vittorio Ferrari
Abstract
We propose a method for annotating videos of complex multi-object scenes with a globally-consistent 3D representation of the objects. We annotate each object with a CAD model from a database, and place it in the 3D coordinate frame of the scene with a 9-DoF pose transformation. Our method is semi-automatic and works on commonly-available RGB videos, without requiring a depth sensor. Many steps are performed automatically, and the tasks performed by humans are simple, well-specified, and require only limited reasoning in 3D. This makes them feasible for crowd-sourcing and has allowed us to construct a large-scale dataset by annotating real-estate videos from YouTube. Our dataset CAD-Estate offers 101k instances of 12k unique CAD models placed in the 3D representations of 20k videos. In comparison to Scan2CAD, the largest existing dataset with CAD model annotations on real scenes, CAD-Estate has 7× more instances and 4× more unique CAD models. We showcase the benefits of pre-training a Mask2CAD model on CAD-Estate for the task of automatic 3D object reconstruction and pose estimation, demonstrating that it leads to performance improvements on the popular Scan2CAD benchmark. The dataset is available at https://github.com/google-research/cad-estate.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e39c251-86b7-4342-bcd6-8a122cc014feCited by top-tier papers6
- SAM 3D: 3Dfy Anything in ImagesXingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang et al.CVPR 2026 · 280 citations
- LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D ScansZhening Huang, Xiaoyang Wu, Fangcheng Zhong, Hengshuang Zhao et al.NeurIPS 2025 · 26 citations
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single ImageZe-Xin Yin, Liu Liu, Xinjie wang, Wei Sui et al.CVPR 2026 · 10 citations
- Diorama: Unleashing Zero-Shot Single-View 3D Indoor Scene ModelingQirui Wu, Denys Iliash, Daniel Ritchie, Manolis Savva et al.ICCV 2025 · 4 citations
- LASA: Instance Reconstruction from Real Scans using A Large-scale Aligned Shape Annotation DatasetHaolin Liu, Chongjie Ye, Yinyu Nie, Yingfan He et al.CVPR 2024 · 1 citation
Builds on10
- Tracking Without Bells and WhistlesPhilipp Bergmann, Tim Meinhardt, Laura Leal-TaixéICCV 2019 · 1,030 citations
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone et al.ICCV 2021 · 686 citations
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar et al.ICCV 2021 · 633 citations
- 3D-FRONT: 3D Furnished Rooms with layOuts and semaNTicsHuan Fu, Bowen Cai, Lin Gao, Lingxiao Zhang et al.ICCV 2021 · 419 citations
- ABO: Dataset and Benchmarks for Real-World 3D Object UnderstandingJasmine Collins, Shubham Goel, Kenan Deng, Achleshwar Luthra et al.CVPR 2022 · 117 citations
Related papers
- RGBD Objects in the Wild: Scaling Real-World 3D Object Learning from RGB-D VideosHongchi Xia, Yang Fu, Sifei Liu, Xiaolong WangCVPR 2024 · 14 citations
- MultiScan: Scalable RGBD scanning for 3D environments with articulated objectsYongsen Mao, Yiming Zhang, Hanxiao Jiang, Angel X. Chang et al.NeurIPS 2022 · 84 citations
- Learning Local RGB-to-CAD Correspondences for Object Pose EstimationGeorgios Georgakis, Srikrishna Karanam, Ziyan Wu, Jana KoseckaICCV 2019 · 25 citations
- UnCommon Objects in 3DXingchen Liu, Piyush Tayal, Jianyuan Wang, Jesus Zarzar et al.CVPR 2025
- Patch2CAD: Patchwise Embedding Learning for In-the-Wild Shape Retrieval from a Single ImageWeicheng Kuo, Anelia Angelova, Tsung-Yi Lin, Angela DaiICCV 2021 · 42 citations
