Zoo3D: Zero-Shot 3D Object Detection at Scene Level
Andrey Lemeshko, Bulat Gabdullin, Nikita Drozdov, Anton Konushin, Danila Rukhovich, Maksim Kolodiazhnyi
Abstract
3D object detection is fundamental for spatial understanding. Real-world environments demand models capable of recognizing diverse, previously unseen objects, which remains a major limitation of closed-set methods. Existing open-vocabulary 3D detectors relax annotation requirements but still depend on training scenes, either as point clouds or images. We take this a step further by introducing Zoo3D, the first training-free 3D object detection framework. Our method constructs 3D bounding boxes via graph clustering of 2D instance masks, then assigns semantic labels using a novel open-vocabulary module with best-view selection and view-consensus mask generation. Zoo3D operates in two modes: the zero-shot Zoo3D, which requires no training at all, and the self-supervised Zoo3D, which refines 3D box prediction by training a class-agnostic detector on Zoo3D-generated pseudo labels. Furthermore, we extend Zoo3D beyond point clouds to work directly with posed and even unposed images. Across ScanNet200 and ARKitScenes benchmarks, both Zoo3D and Zoo3D achieve state-of-the-art results in open-vocabulary 3D object detection. Remarkably, our zero-shot Zoo3D outperforms all existing self-supervised methods, hence demonstrating the power and adaptability of training-free, off-the-shelf approaches for real-world 3D understanding. Code is available at https://github.com/col14m/zoo3d .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39bc3547-9cda-4c08-b076-e7e291591efcBuilds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Deep Hough Voting for 3D Object Detection in Point CloudsCharles R. Qi, Or Litany, Kaiming He, Leonidas J. GuibasICCV 2019 · 1,467 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
Related papers
- OpenM3D: Open Vocabulary Multi-View Indoor 3D Object Detection without Human AnnotationsPeng-Hao Hsu, Ke Zhang, Fu-En Wang, Tao Tu et al.ICCV 2025 · 3 citations
- Towards 3D Objectness Learning in an Open WorldTaichi Liu, Zhenyu Wang, Ruofeng Liu, Guang Wang et al.NeurIPS 2025 · 2 citations
- Training an Open-Vocabulary Monocular 3D Detection Model without 3D DataRui Huang, Henry Zheng, Yan Wang, Zhuofan Xia et al.NeurIPS 2024 · 26 citations
- OpenMask3D: Open-Vocabulary 3D Instance SegmentationAyça Takmaz, Elisabetta Fedele, Robert W. Sumner, Marc Pollefeys et al.NeurIPS 2023 · 389 citations
- OpenBox: Annotate Any Bounding Boxes in 3DIn-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio et al.NeurIPS 2025 · 7 citations
