Detect Anything 3D in the Wild
Hanxue Zhang, Haoran Jiang, Qingsong Yao, Yanan Sun, Renrui Zhang, Hao Zhao, Hongyang Li, Hongzi Zhu, Zetong Yang
Abstract
Despite the success of deep learning in close-set 3D object detection, existing approaches struggle with zero-shot generalization to novel objects and camera configurations. We introduce DetAny3D, a promptable 3D detection foundation model capable of detecting any novel object under arbitrary camera configurations using only monocular inputs. Training a foundation model for 3D detection is fundamentally constrained by the limited availability of annotated 3D data, which motivates DetAny3D to leverage the rich prior knowledge embedded in extensively pre-trained 2D foundation models to compensate for this scarcity. To effectively transfer 2D knowledge to 3D, DetAny3D incorporates two core modules: the 2D Aggregator, which aligns features from different foundation models, and the Interpreter with Zero-Embedding Mapping, which stabilizes early training in 2D-to-3D knowledge transfer. Experimental results validate the strong generalization of our DetAny3D, which not only achieves state-of-the-art performance on unseen categories and novel camera configurations, but also surpasses most competitors on in-domain data. DetAny3D sheds light on the potential of the 3D foundation model for diverse applications in real-world scenarios, e.g., rare object detection in autonomous driving, and demonstrates promise for further exploration of 3D-centric tasks in open-world settings. More visualization results can be found at our code repository.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- SpatialScore: Towards Comprehensive Evaluation for Spatial IntelligenceHaoning Wu, Xiao Huang, Yaohui Chen, Ya Zhang et al.CVPR 2026 · 12 citations
- DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid ThinkingWeicheng Zheng, Xiaofei Mao, Nanfei Ye, Pengxiang Li et al.ICLR 2026 · 8 citations
- LabelAny3D: Label Any Object 3D in the WildJin Yao, Radowan Mahmud Redoy, Sebastian G. Elbaum, Matthew Dwyer et al.NeurIPS 2025 · 8 citations
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-SightYunze Man, Shihao Wang, Guowen Zhang, Johan Bjorck et al.CVPR 2026 · 6 citations
- Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object DetectionXiaojian Lin, Wenxin Zhang, Yuchu Jiang, Wangyu Wu et al.ACM MM 2025 · 3 citations
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion FrameworkTingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia et al.NeurIPS 2022 · 762 citations
- Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene UnderstandingMike Roberts, Jason Ramapuram, Anurag Ranjan, Atulit Kumar et al.ICCV 2021 · 633 citations
Related papers
- Towards 3D Objectness Learning in an Open WorldTaichi Liu, Zhenyu Wang, Ruofeng Liu, Guang Wang et al.NeurIPS 2025 · 2 citations
- Probing the 3D Awareness of Visual Foundation ModelsMohamed El Banani, Amit Raj, Kevis-Kokitsi Maninis, Abhishek Kar et al.CVPR 2024
- Street Gaussians Without 3D Object TrackerRuida Zhang, Chengxi Li, Chenyangguang Zhang, Xingyu Liu et al.ICCV 2025 · 1 citation
- Back to 3D: Few-Shot 3D Keypoint Detection with Back-Projected 2D FeaturesThomas Wimmer, Peter Wonka, Maks OvsjanikovCVPR 2024
- General Object Foundation Model for Images and Videos at ScaleJunfeng Wu, Yi Jiang, Qihao Liu, Zehuan Yuan et al.CVPR 2024 · 36 citations
