Geometry of Lightning Self-Attention: Identifiability and Dimension
Nathan W. Henry, Giovanni Luca Marchetti, Kathlén Kohn
Abstract
We consider function spaces defined by self-attention networks without normalization, and theoretically analyze their geometry. Since these networks are polynomial, we rely on tools from algebraic geometry. In particular, we study the identifiability of deep attention by providing a description of the generic fibers of the parametrization for an arbitrary number of layers and, as a consequence, compute the dimension of the function space. Additionally, for a single-layer model, we characterize the singular and boundary points. Finally, we formulate a conjectural extension of our results to normalized self-attention networks, prove it for a single layer, and numerically verify it in the deep case. Figure 1: A slice of the space of lightning self-attention mechanisms. *Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Identifiability of Deep Polynomial Neural NetworksKonstantin Usevich, Ricardo Augusto Borsoi, Clara Dérand, Marianne ClauselNeurIPS 2025 · 21 citations
- Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural NetworksVahid Shahverdi, Giovanni Luca Marchetti, Kathlén KohnICLR 2026 · 11 citations
- Topology and geometry of the learning space of ReLU networks: connectivity and singularitiesMarco Nurisso, Pierrick Leroy, Giovanni Petri, Francesco VaccarinoICLR 2026 · 6 citations
- Identifiable Equivariant Networks are Layerwise EquivariantVahid Shahverdi, Giovanni Luca Marchetti, Georg Bökman, Kathlén KohnICML 2026
Builds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 394 citations
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning DynamicsDaniel Kunin, Javier Sagastuy-Breña, Surya Ganguli, Daniel L. K. Yamins et al.ICLR 2021 · 100 citations
- Hidden Symmetries of ReLU NetworksJ. Elisenda Grigsby, Kathryn Lindsey, David RolnickICML 2023 · 35 citations
- Symmetry Teleportation for Accelerated OptimizationBo Zhao, Nima Dehmamy, Robin Walters, Rose YuNeurIPS 2022 · 33 citations
Related papers
- Attention Mechanism, Max-Affine Partition, and Universal ApproximationHude Liu, Jerry Yao-Chieh Hu, Zhao Song, Han LiuNeurIPS 2025 · 12 citations
- Lipschitz normalization for self-attention layers with application to graph neural networksGeorge Dasoulas, Kevin Scaman, Aladin VirmauxICML 2021 · 55 citations
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 172 citations
- Spurious Valleys and Clustering Behavior of Neural NetworksSamuele PollaciICML 2023 · 1 citation
- Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor ProgramsMana Sakai, Ryo Karakida, Masaaki ImaizumiNeurIPS 2025 · 5 citations
