Unveiling the Visual Counting Bottleneck in Vision-Language Models
Xingzhou Pang, Yifan Hou, Junling Wang, Mrinmaya Sachan
Abstract
While Large Vision-Language Models (VLMs) excel at interpolation, they suffer catastrophic failures in systematic generalization, most notably in visual counting beyond training distributions. In this work, we investigate this extrapolation bottleneck by deconstructing visual counting into three cognitive stages: object individuation, abstract magnitude representation, and symbolic decoding. Using a controlled environment of synthetic Go game boards, we isolate the specific mechanism of failure. Contrary to the hypothesis that models suffer from perceptual errors, we demonstrate via linear probing that visual backbones maintain robust, linearly separable representations of quantity well into the extrapolation regime. Furthermore, models retain latent magnitude awareness, successfully performing comparative reasoning on quantities they fail to enumerate. We pinpoint the collapse to the Symbolic Decoding stage, where the model fails to project valid visual magnitudes onto discrete tokens. Our findings support a Fractured Magnitude Hypothesis: VLMs fail to acquire a Universal Number Space, instead learning disjoint, modality-specific statistical manifolds that prevent cross-modal grounding for unseen pairings. We validate our findings on the state-of-the-art foundation model, suggesting that bridging the extrapolation gap requires inductive priors that enforce unified magnitude representations rather than simply scaling training data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4c574540-2bf5-40a2-b4c7-5101646ecdbdBuilds on16
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- The Impact of Positional Encoding on Length Generalization in TransformersAmirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das et al.NeurIPS 2023 · 444 citations
Related papers
- Understanding the Limits of Vision Language Models Through the Lens of the Binding ProblemDeclan Campbell, Sunayana Rane, Tyler Giallanza, Nicolò De Sabbata et al.NeurIPS 2024 · 101 citations
- Understanding Counting Mechanisms in Large Language and Vision-Language ModelsHosein Hasani, Amirmohammad Izadi, Fatemeh Askari, Mobin Bagherian et al.CVPR 2026 · 5 citations
- Seeing Without Understanding: Disentangling Perception, Reasoning, and Simulation in VLM GameplayDingyang Jin, Jiawei He, Calvin Lo, Steven Hu et al.ICML 2026
- Caption This, Reason That: VLMs Caught in the MiddleZihan Weng, Lucas Gomez, Taylor W. Webb, Pouya BashivanNeurIPS 2025 · 3 citations
- Response Wide Shut? Surprising Observations in Basic Vision Language Model CapabilitiesShivam Chandhok, Wan-Cyuan Fan, Vered Shwartz, Vineeth N. Balasubramanian et al.ACL 2025
