Hyperbolic Learning with Synthetic Captions for Open-World Detection
Fanjie Kong, Yanbei Chen, Jiarui Cai, Davide Modolo
Abstract
Open-world detection poses significant challenges, as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use large-scale manual annotated caption datasets for training, which are extremely expensive to collect. Instead, we propose to transfer knowledge from vision-language models (VLMs) to enrich the open-vocabulary descriptions automatically. Specifically, we bootstrap dense synthetic captions using pre-trained VLMs to provide rich descriptions on different regions in images, and incorporate these captions to train a novel detector that generalizes to novel concepts. To mitigate the noise caused by hallucination in synthetic captions, we also propose a novel hyperbolic visionlanguage learning approach to impose a hierarchy between visual and caption embeddings. We call our detector "Hy-perLearner". We conduct extensive experiments on a wide variety of open-world detection benchmarks (COCO, LVIS, Object Detection in the Wild, RefCOCO) and our results show that our model consistently outperforms existing stateof-the-art methods, such as GLIP, GLIPv2 and Grounding DINO, when using the same backbone.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bcfd76a-adca-4fa2-bcbf-837b91486f20Cited by top-tier papers9
- Learning Fine-grained Domain Generalization via Hyperbolic State Space HallucinationQi Bi, Jingjun Yi, Haolan Zhan, Wei Ji et al.AAAI 2025 · 8 citations
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language ModelsZelin Peng, Zhengqin Xu, Qingyang Liu, Xiaokang Yang et al.NeurIPS 2025 · 5 citations
- Learning Visual Hierarchies in Hyperbolic Space for Image RetrievalZiwei Wang, Sameera Ramasinghe, Chenchen Hu, Julien Monteil et al.ICCV 2025 · 4 citations
- Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional VideosSagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan et al.CVPR 2025
- A Hybrid Space Model for Misaligned Multi-modality Image FusionYi Xiao, Jia Wang, Zhu Liu, Di Wang et al.AAAI 2026
Builds on35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- CapDet: Unifying Dense Captioning and Open-World Detection PretrainingYanxin Long, Youpeng Wen, Jianhua Han, Hang Xu et al.CVPR 2023
- Learning to Prompt for Open-Vocabulary Object Detection with Vision-Language ModelYu Du, Fangyun Wei, Zihe Zhang, Miaojing Shi et al.CVPR 2022 · 311 citations
- Learning Object-Language Alignments for Open-Vocabulary Object DetectionChuang Lin, Peize Sun, Yi Jiang, Ping Luo et al.ICLR 2023 · 36 citations
- Retrieval-Augmented Open-Vocabulary Object DetectionJooyeon Kim, Eulrang Cho, Sehyung Kim, Hyunwoo J. KimCVPR 2024
- Towards Universal Perception through Language-Guided Open-World Object DetectionZihan Wang, Yunhang Shen, Yuan Fang, Zuwei Long et al.ACM MM 2025 · 1 citation
