On the Intrinsic Limits of Transformer Image Embeddings in Non-Solvable Spatial Reasoning
Siyi Lyu, Quan Liu, Feng Yan
Abstract
Vision Transformers (ViTs) excel in semantic recognition but exhibit systematic failures in spatial reasoning tasks such as mental rotation. While often attributed to data scale, this work argues that the limitation arises from the intrinsic circuit complexity of the architecture. By formalizing spatial understanding as a Group Homomorphism Problem—where latent embeddings preserve the algebraic structure of physical transformations acting on images—we identify a fundamental computational bottleneck. Specifically, for non-solvable groups (e.g., ), maintaining such structure-preserving embeddings is lower-bounded by the Word Problem, which is -complete. In contrast, constant-depth ViTs with polynomial precision are strictly bounded by the complexity class . Under the standard conjecture , a complexity boundary emerges: constant-depth architectures lack the logical depth required to capture non-solvable spatial structures in a single forward pass. To empirically validate this theoretical gap, we propose the Latent Space Algebra (LSA) benchmark, which reveals a significant degradation in ViT representations as the compositional depth of non-solvable tasks increases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c53073ce-be88-4908-ba8f-e1bef145ca80Builds on7
- The Illusion of State in State-Space ModelsWilliam Merrill, Jackson Petty, Ashish SabharwalICML 2024 · 157 citations
- A Logic for Expressing Log-Precision TransformersWilliam Merrill, Ashish SabharwalNeurIPS 2023 · 87 citations
- Group Equivariant Stand-Alone Self-Attention For VisionDavid W. Romero, Jean-Baptiste CordonnierICLR 2021 · 72 citations
- Circuit Complexity Bounds for RoPE-based Transformer ArchitectureBo Chen, Xiaoyu Li, Yingyu Liang, Jiangxuan Long et al.EMNLP 2025 · 33 citations
- Beyond the Doors of Perception: Vision Transformers Represent Relations Between ObjectsMichael A. Lepori, Alexa R. Tartaglini, Wai Keen Vong, Thomas Serre et al.NeurIPS 2024 · 22 citations
Related papers
- 3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language ModelsShaoxiong Zhan, Yanlin Lai, Zheng Liu, Zijian Lin et al.ICML 2026 · 6 citations
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language ModelsMengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang et al.ICLR 2026 · 109 citations
- SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMsYuyou Zhang, Radu Corcodel, Chiori Hori, Anoop Cherian et al.ICLR 2026 · 11 citations
- Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed SpacesChen Yang, Guanxin Lin, Youquan He, Peiyao Chen et al.ICML 2026
- Defining and Evaluating Visual Language Models' Basic Spatial Abilities: A Perspective from PsychometricsWenrui Xu, Dalin Lyu, Weihang Wang, Jie Feng et al.ACL 2025 · 19 citations
