VoxDet: Voxel Learning for Novel Instance Detection
Bowen Li, Jiashun Wang, Yaoyu Hu, Chen Wang, Sebastian A. Scherer
Abstract
Detecting unseen instances based on multi-view templates is a challenging problem due to its open-world nature. Traditional methodologies, which primarily rely on 2D representations and matching techniques, are often inadequate in handling pose variations and occlusions. To solve this, we introduce VoxDet, a pioneer 3D geometry-aware framework that fully utilizes the strong 3D voxel representation and reliable voxel matching mechanism. VoxDet first ingeniously proposes template voxel aggregation (TVA) module, effectively transforming multi-view 2D images into 3D voxel features. By leveraging associated camera poses, these features are aggregated into a compact 3D template voxel. In novel instance detection, this voxel representation demonstrates heightened resilience to occlusion and pose variations. We also discover that a 3D reconstruction objective helps to pre-train the 2D-3D mapping in TVA. Second, to quickly align with the template voxel, VoxDet incorporates a Query Voxel Matching (QVM) module. The 2D queries are first converted into their voxel representation with the learned 2D-3D mapping. We find that since the 3D voxel representations encode the geometry, we can first estimate the relative rotation and then compare the aligned voxels, leading to improved accuracy and efficiency. In addition to method, we also introduce the first instance detection benchmark, RoboTools, where 20 unique instances are video-recorded with camera extrinsic. RoboTools also provides 24 challenging cluttered scenarios with more than 9k box annotations. Exhaustive experiments are conducted on the demanding LineMod-Occlusion, YCB-video, and RoboTools benchmarks, where VoxDet outperforms various 2D baselines remarkably with faster speed. To the best of our knowledge, VoxDet is the first to incorporate implicit 3D knowledge for 2D novel instance detection tasks. Our code, data, raw results, and pre-trained models are public at https://github.com/Jaraxxus-Me/VoxDet .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9713ccc-eff6-499d-969d-194b613a2b29Cited by top-tier papers3
- VoxDet: Rethinking 3D Semantic Scene Completion as Dense Object DetectionWuyang Li, Zhu Yu, Alexandre AlahiNeurIPS 2025 · 3 citations
- Find your Needle: Small Object Image Retrieval via Multi-Object Attention OptimizationMichael Green, Matan Levy, Issar Tzachor, Dvir Samuel et al.NeurIPS 2025 · 1 citation
- Solving Instance Detection from an Open-World PerspectiveQianqian Shen, Yunhan Zhao, Nahyun Kwon, Jeeeun Kim et al.CVPR 2025
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- MixFormer: End-to-End Tracking with Iterative Mixed AttentionYutao Cui, Cheng Jiang, Limin Wang, Gangshan WuCVPR 2022 · 746 citations
- Frustratingly Simple Few-Shot Object DetectionXin Wang, Thomas E. Huang, Joseph Gonzalez, Trevor Darrell et al.ICML 2020 · 723 citations
Related papers
- Viewpoint Equivariance for Multi-View 3D Object DetectionDian Chen, Jie Li, Vitor Guizilini, Rares Ambrus et al.CVPR 2023
- Boosting Multi-View Indoor 3D Object Detection Via Adaptive 3D Volume ConstructionRunmin Zhang, Zhu Yu, Si-Yuan Cao, Lingyu Zhu et al.ICCV 2025 · 3 citations
- CoDA: Collaborative Novel Box Discovery and Cross-modal Alignment for Open-vocabulary 3D Object DetectionYang Cao, Yihan Zeng, Hang Xu, Dan XuNeurIPS 2023 · 69 citations
- ImGeoNet: Image-induced Geometry-aware Voxel Representation for Multi-view 3D Object DetectionTao Tu, Shun-Po Chuang, Yu-Lun Liu, Cheng Sun et al.ICCV 2023 · 16 citations
- Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D Mask GuidancePhuc D. A. Nguyen, Tuan Duc Ngo, Evangelos Kalogerakis, Chuang Gan et al.CVPR 2024 · 45 citations
