The Devil is in the Fine-Grained Details: Evaluating open-Vocabulary Object Detectors for Fine-Grained Understanding
Lorenzo Bianchi, Fabio Carrara, Nicola Messina, Claudio Gennaro, Fabrizio Falchi
摘要
Recent advancements in large vision-language models enabled visual object detection in open-vocabulary scenarios, where object classes are defined in free-text formats during inference. In this paper, we aim to probe the stateof-the-art methods for open-vocabulary object detection to determine to what extent they understand fine-grained properties of objects and their parts. To this end, we introduce an evaluation protocol based on dynamic vocabulary generation to test whether models detect, discern, and assign the correct fine-grained description to objects in the presence of hard-negative classes. We contribute with a benchmark suite of increasing difficulty and probing different properties like color, pattern, and material. We further enhance our investigation by evaluating several stateof-the-art open-vocabulary object detectors using the proposed protocol and find that most existing solutions, which shine in standard open-vocabulary benchmarks, struggle to accurately capture and distinguish finer object details. We conclude the paper by highlighting the limitations of current methodologies and exploring promising research directions to overcome the discovered drawbacks. Data and code are available at https://lorebianchi98.github.io/FG-OVD/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment ModelChunyu Xie, Bin Wang, Fanjing Kong, Jincheng Li 等ICML 2026 · 被引用 14 次
- Where, What, Why: Towards Explainable Driver Attention PredictionYuchen Zhou, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin 等ICCV 2025 · 被引用 8 次
- GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object DetectionJiaming Li, Zhijia Liang, Weikai Chen, Lin Ma 等NeurIPS 2025 · 被引用 6 次
- β-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language AlignmentFatimah Zohra, Chen Zhao, Hani Itani, Bernard GhanemCVPR 2026 · 被引用 6 次
- OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene UnderstandingYoujun Zhao, Jiaying Lin, Shuquan Ye, Qianshi Pang 等AAAI 2026 · 被引用 5 次
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve 等ICCV 2021 · 被引用 1,114 次
- Scaling Open-Vocabulary Object DetectionMatthias Minderer, Alexey A. Gritsenko, Neil HoulsbyNeurIPS 2023 · 被引用 482 次
相关 Paper
- Open-vocabulary Attribute DetectionMaría Alejandra Bravo, Sudhanshu Mittal, Simon Ging, Thomas BroxCVPR 2023
- Generating Enhanced Negatives for Training Language-Based Object DetectorsShiyu Zhao, Long Zhao, Vijay Kumar B. G, Yumin Suh 等CVPR 2024 · 被引用 4 次
- How to Evaluate the Generalization of Detection? A Benchmark for Comprehensive Open-Vocabulary DetectionYiyang Yao, Peng Liu, Tiancheng Zhao, Qianqian Zhang 等AAAI 2024 · 被引用 18 次
- OmniLabel: A Challenging Benchmark for Language-Based Object DetectionSamuel Schulter, Vijay Kumar B. G, Yumin Suh, Konstantinos M. Dafnis 等ICCV 2023 · 被引用 18 次
- Stepping Out of Similar Semantic Space for Open-Vocabulary SegmentationYong Liu, Song-Li Wu, Sule Bai, Jiahao Wang 等ICCV 2025 · 被引用 6 次
