GOOD: Exploring geometric cues for detecting objects in an open world
Haiwen Huang, Andreas Geiger, Dan Zhang
Abstract
We address the task of open-world class-agnostic object detection, i.e., detecting every object in an image by learning from a limited number of base object classes. State-of-the-art RGB-based models suffer from overfitting the training classes and often fail at detecting novel-looking objects. This is because RGB-based models primarily rely on appearance similarity to detect novel objects and are also prone to overfitting short-cut cues such as textures and discriminative parts. To address these shortcomings of RGB-based object detectors, we propose incorporating geometric cues such as depth and normals, predicted by general-purpose monocular estimators. Specifically, we use the geometric cues to train an object proposal network for pseudo-labeling unannotated novel objects in the training set. Our resulting Geometry-guided Open-world Object Detector (GOOD) significantly improves detection recall for novel object categories and already performs well with only a few training classes. Using a single "person" class for training on the COCO dataset, GOOD surpasses SOTA methods by 5.0% AR@100, a relative improvement of 24%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8db30e47-023b-4c51-a0df-69141ce29ee9Cited by top-tier papers4
- AG3D: Learning to Generate 3D Avatars from 2D Image CollectionsZijian Dong, Xu Chen, Jinlong Yang, Michael J. Black et al.ICCV 2023 · 76 citations
- Exploring Transformers for Open-world Instance SegmentationJiannan Wu, Yi Jiang, Bin Yan, Huchuan Lu et al.ICCV 2023 · 7 citations
- Detecting Open World Objects via Partial Attribute AssignmentMuli Yang, Gabriel James Goenawan, Huaiyuan Qin, Kai Han et al.CVPR 2025
- v-CLR: View-Consistent Learning for Open-World Instance SegmentationChang-Bin Zhang, Jinhong Ni, Yujie Zhong, Kai HanCVPR 2025
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- MonoSDF: Exploring Monocular Geometric Cues for Neural Implicit Surface ReconstructionZehao Yu, Songyou Peng, Michael Niemeyer, Torsten Sattler et al.NeurIPS 2022 · 670 citations
Related papers
- Towards 3D Objectness Learning in an Open WorldTaichi Liu, Zhenyu Wang, Ruofeng Liu, Guang Wang et al.NeurIPS 2025 · 2 citations
- OW-DETR: Open-world Detection TransformerAkshita Gupta, Sanath Narayan, K. J. Joseph, Salman Khan et al.CVPR 2022 · 209 citations
- 3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object DetectionYung-Hsu Yang, Luigi Piccinelli, Mattia Segù, Siyuan Li et al.ICCV 2025 · 2 citations
- Training an Open-Vocabulary Monocular 3D Detection Model without 3D DataRui Huang, Henry Zheng, Yan Wang, Zhuofan Xia et al.NeurIPS 2024 · 26 citations
- PROB: Probabilistic Objectness for Open World Object DetectionOrr Zohar, Kuan-Chieh Wang, Serena YeungCVPR 2023
