On Large Multimodal Models as Open-World Image Classifiers
Alessandro Conti, Massimiliano Mancini, Enrico Fini, Yiming Wang, Paolo Rota, Elisa Ricci
摘要
Traditional image classification requires a predefined list of semantic categories. In contrast, Large Multimodal Models (LMMs) can sidestep this requirement by classifying images directly using natural language (e.g., answering the prompt "What is the main object in the image?"). Despite this remarkable capability, most existing studies on LMM classification performance are surprisingly limited in scope, often assuming a closed-world setting with a predefined set of categories. In this work, we address this gap by thoroughly evaluating LMM classification performance in a truly open-world setting. We first formalize the task and introduce an evaluation protocol, defining various metrics to assess the alignment between predicted and ground truth classes. We then evaluate 13 models across 10 benchmarks, encompassing prototypical, non-prototypical, finegrained, and very fine-grained classes, demonstrating the challenges LMMs face in this task. Further analyses based on the proposed metrics reveal the types of errors LMMs make, highlighting challenges related to granularity and fine-grained capabilities, showing how tailored prompting and reasoning can alleviate them. Code is available at https://github.com/altndrr/lmms-owc.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual RecognitionYuwen Tan, Yuan Qing, Boqing GongCVPR 2026 · 被引用 6 次
- On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action UnderstandingZhanzhong Pang, Dibyadip Chatterjee, Fadime Sener, Angela YaoICLR 2026 · 被引用 1 次
- Specificity-aware reinforcement learning for fine-grained open-world classificationSamuele Angheben, Davide Berasi, Alessandro Conti, Elisa Ricci 等CVPR 2026
- Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal ModelsHulingxiao He, Zhi Tan, Yuxin PengICML 2026
它引用的顶会 Paper32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive EvaluationHong-Tao Yu, Yuxin Peng, Serge J. Belongie, Xiu-Shen WeiICLR 2026 · 被引用 21 次
- Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMsDmitry Demidov, Muhammad Zaigham Zaheer, Zongyan Han, Omkar Thawakar 等CVPR 2026
- What does a platypus look like? Generating customized prompts for zero-shot image classificationSarah M. Pratt, Ian Covert, Rosanne Liu, Ali FarhadiICCV 2023 · 被引用 343 次
- VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language ModelsMingjie Xu, Jinpeng Chen, Yuzhi Zhao, Jason Chun Lok Li 等AAAI 2026
- Multi-Modal Classifiers for Open-Vocabulary Object DetectionPrannay Kaul, Weidi Xie, Andrew ZissermanICML 2023 · 被引用 69 次
