Cubify Anything: Scaling Indoor 3D Object Detection
Justin Lazarow, David Griffiths, Gefen Kohavi, Francisco Crespo, Afshin Dehghan
2025Year
15Top-tier citations
Abstract
ScanNet v2 ARKitScenes . CA-1M is the first dataset to provide explicit 3D boxes which cover the full richness of objects while being both spatially accurate and pixel-perfect with respect to each frame. Existing datasets like SUN RGB-D, ScanNet v2, ARKitScenes are either small, coarsely labeled, or lack accurate mappings from world to image space. Since ARKitScenes and CA-1M are labeled on the same underlying data, we can show the effect of exhaustive labeling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han et al.NeurIPS 2025 · 159 citations
- Scaling Spatial Intelligence with Multimodal Foundation ModelsZhongang Cai, Wang Ruisi, Chenyang Gu, Fanyi Pu et al.CVPR 2026 · 81 citations
- LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D ScansZhening Huang, Xiaoyang Wu, Fangcheng Zhong, Hengshuang Zhao et al.NeurIPS 2025 · 26 citations
- Vlaser: Vision-Language-Action Model with Synergistic Embodied ReasoningGanlin Yang, Tianyi Zhang, Haoran Hao, Weiyun Wang et al.ICLR 2026 · 23 citations
- SpatialScore: Towards Comprehensive Evaluation for Spatial IntelligenceHaoning Wu, Xiao Huang, Yaohui Chen, Ya Zhang et al.CVPR 2026 · 12 citations
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
Related papers
- ARKit LabelMaker: A New Scale for Indoor 3D Scene UnderstandingGuangda Ji, Silvan Weder, Francis Engelmann, Marc Pollefeys et al.CVPR 2025
- Rooms from Motion: Un-posed Indoor 3D Object Detection as Localization and MappingJustin Lazarow, Kai Kang, Afshin DehghanNeurIPS 2025 · 2 citations
- Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume ConstructionRunmin Zhang, Zhu Yu, Si-Yuan Cao, Lingyu Zhu et al.ICCV 2025 · 3 citations
- Zoo3D: Zero-Shot 3D Object Detection at Scene LevelAndrey Lemeshko, Bulat Gabdullin, Nikita Drozdov, Anton Konushin et al.CVPR 2026 · 5 citations
- OpenM3D: Open Vocabulary Multi-View Indoor 3D Object Detection without Human AnnotationsPeng-Hao Hsu, Ke Zhang, Fu-En Wang, Tao Tu et al.ICCV 2025 · 3 citations
