VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
Shmuel Berman, Jia Deng
Abstract
Vision-Language Models (VLMs) excel at complex visual tasks such as VQA and chart understanding, yet recent work suggests they struggle with simple perceptual tests. We present an evaluation that tests vision-language models' capacity for nonlocal visual reasoning-reasoning that requires chaining evidence collected from multiple, possibly distant, regions of an image. We isolate three distinct forms of nonlocal vision: comparative perception, which demands holding two images in working memory and comparing them; saccadic search, which requires making discrete, evidence-driven jumps to locate successive targets; and smooth visual search, which involves searching smoothly along a continuous contour. Flagship models (e.g. GPT-5, Gemini 2.5 Pro, Claude Sonnet 4), even those that perform well on prior primitive-vision benchmarks, fail these tests and barely exceed random accuracy on two variants of our tasks that are trivial for humans. Our structured evaluation suite allows us to test if VLMs can perform similar visual algorithms to humans. Our findings show that despite gains in raw visual acuity, current models lack core visual reasoning capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75917bd8-a7fa-44fa-98a8-ea7d6369d775Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingKenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu et al.ICML 2023 · 426 citations
- UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and ReasoningAhmed Masry, Parsa Kavehzadeh, Do Xuan Long, Enamul Hoque et al.EMNLP 2023 · 48 citations
- Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language ModelsMatt Deitke, Christopher Clark, Sangho Lee, Rohun Tripathi et al.CVPR 2025
- Hallusionbench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language ModelsTianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian et al.CVPR 2024
Related papers
- Caption This, Reason That: VLMs Caught in the MiddleZihan Weng, Lucas Gomez, Taylor W. Webb, Pouya BashivanNeurIPS 2025 · 3 citations
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para et al.CVPR 2026 · 3 citations
- Bongard in Wonderland: Visual Puzzles that Still Make AI Go Mad?Antonia Wüst, Tim Nelson Tobiasch, Lukas Helff, Inga Ibs et al.ICML 2025
- Charts-of-Thought: Enhancing LLM Visualization Literacy Through Structured Data ExtractionAmit Kumar Das, Mohammad Tarun, Klaus MuellerIEEE VIS 2025 · 6 citations
- Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable EventsAditya Chinchure, Sahithya Ravi, Raymond T. Ng, Vered Shwartz et al.CVPR 2025
