On the Intrinsic Limits of Transformer Image Embeddings in Non-Solvable Spatial Reasoning
Siyi Lyu, Quan Liu, Feng Yan
摘要
Vision Transformers (ViTs) excel in semantic recognition but exhibit systematic failures in spatial reasoning tasks such as mental rotation. While often attributed to data scale, this work argues that the limitation arises from the intrinsic circuit complexity of the architecture. By formalizing spatial understanding as a Group Homomorphism Problem—where latent embeddings preserve the algebraic structure of physical transformations acting on images—we identify a fundamental computational bottleneck. Specifically, for non-solvable groups (e.g., ), maintaining such structure-preserving embeddings is lower-bounded by the Word Problem, which is -complete. In contrast, constant-depth ViTs with polynomial precision are strictly bounded by the complexity class . Under the standard conjecture , a complexity boundary emerges: constant-depth architectures lack the logical depth required to capture non-solvable spatial structures in a single forward pass. To empirically validate this theoretical gap, we propose the Latent Space Algebra (LSA) benchmark, which reveals a significant degradation in ViT representations as the compositional depth of non-solvable tasks increases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- The Illusion of State in State-Space ModelsWilliam Merrill, Jackson Petty, Ashish SabharwalICML 2024 · 被引用 157 次
- A Logic for Expressing Log-Precision TransformersWilliam Merrill, Ashish SabharwalNeurIPS 2023 · 被引用 87 次
- Group Equivariant Stand-Alone Self-Attention For VisionDavid W. Romero, Jean-Baptiste CordonnierICLR 2021 · 被引用 72 次
- Circuit Complexity Bounds for RoPE-based Transformer ArchitectureBo Chen, Xiaoyu Li, Yingyu Liang, Jiangxuan Long 等EMNLP 2025 · 被引用 33 次
- Beyond the Doors of Perception: Vision Transformers Represent Relations Between ObjectsMichael A. Lepori, Alexa R. Tartaglini, Wai Keen Vong, Thomas Serre 等NeurIPS 2024 · 被引用 22 次
相关 Paper
- 3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language ModelsShaoxiong Zhan, Yanlin Lai, Zheng Liu, Zijian Lin 等ICML 2026 · 被引用 6 次
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language ModelsMengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang 等ICLR 2026 · 被引用 109 次
- SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMsYuyou Zhang, Radu Corcodel, Chiori Hori, Anoop Cherian 等ICLR 2026 · 被引用 11 次
- Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed SpacesChen Yang, Guanxin Lin, Youquan He, Peiyao Chen 等ICML 2026
- Defining and Evaluating Visual Language Models' Basic Spatial Abilities: A Perspective from PsychometricsWenrui Xu, Dalin Lyu, Weihang Wang, Jie Feng 等ACL 2025 · 被引用 19 次
