The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
Yuwen Tan, Yuan Qing, Boqing Gong
Abstract
This paper reveals that many open-source large language models (LLMs) lack hierarchical knowledge about our visual world, unaware of even well-established biology taxonomies. This shortcoming makes LLMs a bottleneck for vision LLMs' hierarchical visual recognition (e.g., recognizing Anemone Fish but not Vertebrate). We arrive at these findings using about one million four-choice visual question answering (VQA) tasks constructed from six taxonomies and four image datasets. Interestingly, finetuning a vision LLM using our VQA tasks reaffirms LLMs' bottleneck effect because the VQA tasks improve the LLMs' hierarchical consistency in text-only tasks more than the vision LLMs'. We believe that one cannot make vision LLMs understand our visual world hierarchically until LLMs possess corresponding taxonomy knowledge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 68e2a60a-c37a-4abc-a7ab-d421584127f6Cited by top-tier papers6
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter ItYulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli et al.NeurIPS 2025 · 11 citations
- Free-Grained Hierarchical Visual RecognitionSeulki Park, Zilin Wang, Stella X. YuCVPR 2026 · 3 citations
- Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal ModelsHulingxiao He, Zhi Tan, Yuxin PengCVPR 2026 · 3 citations
- Specificity-aware reinforcement learning for fine-grained open-world classificationSamuele Angheben, Davide Berasi, Alessandro Conti, Elisa Ricci et al.CVPR 2026
- RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMsLogan Lawrence, Oindrila Saha, Rangel Daroya, Mustafa Chasmai et al.CVPR 2026
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- Why are Visually-Grounded Language Models Bad at Image Classification?Yuhui Zhang, Alyssa Unell, Xiaohan Wang, Dhruba Ghosh et al.NeurIPS 2024 · 128 citations
- OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLMYutao Hu, Tianbin Li, Quanfeng Lu, Wenqi Shao et al.CVPR 2024
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language ModelsJiayun Luo, Wan-Cyuan Fan, Lyuyang Wang, Xiangteng He et al.ICLR 2026 · 19 citations
- Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question AnsweringMahsa Mozaffari, Hitesh Sapkota, Xumin Liu, Qi YuICLR 2026
- Vision Language Models are BiasedAn Vo, Khai-Nguyen Nguyen, Mohammad Reza Taesiri, Thi Tuong Vy Dang et al.ICLR 2026 · 68 citations
