The Devil is in the Fine-Grained Details: Evaluating open-Vocabulary Object Detectors for Fine-Grained Understanding
Lorenzo Bianchi, Fabio Carrara, Nicola Messina, Claudio Gennaro, Fabrizio Falchi
Abstract
Recent advancements in large vision-language models enabled visual object detection in open-vocabulary scenarios, where object classes are defined in free-text formats during inference. In this paper, we aim to probe the stateof-the-art methods for open-vocabulary object detection to determine to what extent they understand fine-grained properties of objects and their parts. To this end, we introduce an evaluation protocol based on dynamic vocabulary generation to test whether models detect, discern, and assign the correct fine-grained description to objects in the presence of hard-negative classes. We contribute with a benchmark suite of increasing difficulty and probing different properties like color, pattern, and material. We further enhance our investigation by evaluating several stateof-the-art open-vocabulary object detectors using the proposed protocol and find that most existing solutions, which shine in standard open-vocabulary benchmarks, struggle to accurately capture and distinguish finer object details. We conclude the paper by highlighting the limitations of current methodologies and exploring promising research directions to overcome the discovered drawbacks. Data and code are available at https://lorebianchi98.github.io/FG-OVD/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9d9673d1-5658-4e8c-aace-c9ab1eb2a47fCited by top-tier papers14
- FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment ModelChunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li et al.ICML 2026 · 14 citations
- Where, What, Why: Towards Explainable Driver Attention PredictionYuchen Zhou, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin et al.ICCV 2025 · 8 citations
- GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object DetectionJiaming Li, Zhijia Liang, Weikai Chen, Lin Ma et al.NeurIPS 2025 · 6 citations
- β-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language AlignmentFatimah Zohra, Chen Zhao, Hani Itani, Bernard GhanemCVPR 2026 · 6 citations
- OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene UnderstandingYoujun Zhao, Jiaying Lin, Shuquan Ye, Qianshi Pang et al.AAAI 2026 · 5 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- Scaling Open-Vocabulary Object DetectionMatthias Minderer, Alexey A. Gritsenko, Neil HoulsbyNeurIPS 2023 · 482 citations
Related papers
- Open-vocabulary Attribute DetectionMaría Alejandra Bravo, Sudhanshu Mittal, Simon Ging, Thomas BroxCVPR 2023
- Generating Enhanced Negatives for Training Language-Based Object DetectorsShiyu Zhao, Long Zhao, Vijay Kumar B. G, Yumin Suh et al.CVPR 2024 · 4 citations
- How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary DetectionYiyang Yao, Peng Liu, Tiancheng Zhao, Qianqian Zhang et al.AAAI 2024 · 18 citations
- OmniLabel: A Challenging Benchmark for Language-Based Object DetectionSamuel Schulter, Vijay Kumar B. G, Yumin Suh, Konstantinos M. Dafnis et al.ICCV 2023 · 18 citations
- Stepping Out of Similar Semantic Space for Open-Vocabulary SegmentationYong Liu, Song-Li Wu, Sule Bai, Jiahao Wang et al.ICCV 2025 · 6 citations
