Taxonomy-Aware Evaluation of Vision-Language Models
Vésteinn Snæbjarnarson, Kevin Du, Niklas Stoehr, Serge J. Belongie, Ryan Cotterell, Nico Lang, Stella Frank
Abstract
When a vision-language model (VLM) is prompted to identify an entity depicted in an image, it may answer "I see a conifer," rather than the specific label NORWAY SPRUCE. This raises two issues for evaluation: Firstly, the unconstrained generated text needs to be mapped to the evaluation label space (i.e., CONIFER). Secondly, a useful classification measure should give partial credit to lessspecific, but not incorrect, answers (NORWAY SPRUCE being a type of CONIFER). To meet these requirements, we propose a framework for evaluating unconstrained text predictions such as those generated from a vision-language model against a taxonomy. Specifically, we propose the use of hierarchical precision and recall measures to assess the level of correctness and specificity of predictions with regard to a taxonomy. Experimentally, we first show that existing text similarity measures do not capture taxonomic similarity well. We then develop and compare different methods to map textual VLM predictions onto a taxonomy. This allows us to compute hierarchical similarity measures between the generated text and the ground truth labels. Finally, we analyze modern VLMs on fine-grained visual classification tasks based on our proposed taxonomic evaluation scheme. Data and code are made available at https://github.com/vesteinn/vlm-eval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67940fbe-8653-4ee4-853b-9cf30cf2cd2bCited by top-tier papers6
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter ItYulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli et al.NeurIPS 2025 · 11 citations
- The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual RecognitionYuwen Tan, Yuan Qing, Boqing GongCVPR 2026 · 6 citations
- Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal ModelsHulingxiao He, Zhi Tan, Yuxin PengCVPR 2026 · 3 citations
- Specificity-aware reinforcement learning for fine-grained open-world classificationSamuele Angheben, Davide Berasi, Alessandro Conti, Elisa Ricci et al.CVPR 2026
- RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMsLogan Lawrence, Oindrila Saha, Rangel Daroya, Mustafa Chasmai et al.CVPR 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchySimon Ging, María Alejandra Bravo, Thomas BroxICLR 2024 · 24 citations
- VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic PhenomenaLetitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank et al.ACL 2022 · 147 citations
- Zero-Shot Text-to-Motion Evaluation using Video Language ModelsYuwen Ji, Donglin Wang, Yue ZhangICML 2026
- Trust but Verify: Programmatic VLM Evaluation in the WildViraj Prabhu, Senthil Purushwalkam, An Yan, Caiming Xiong et al.ICCV 2025
- Cross-Modal Taxonomic Generalization in (Vision-) Language ModelsTianyang Xu, Marcelo Sandoval-Castañeda, Karen Livescu, Greg Shakhnarovich et al.ACL 2026
