Geometry of Lightning Self-Attention: Identifiability and Dimension
Nathan W. Henry, Giovanni Luca Marchetti, Kathlén Kohn
摘要
We consider function spaces defined by self-attention networks without normalization, and theoretically analyze their geometry. Since these networks are polynomial, we rely on tools from algebraic geometry. In particular, we study the identifiability of deep attention by providing a description of the generic fibers of the parametrization for an arbitrary number of layers and, as a consequence, compute the dimension of the function space. Additionally, for a single-layer model, we characterize the singular and boundary points. Finally, we formulate a conjectural extension of our results to normalized self-attention networks, prove it for a single layer, and numerically verify it in the deep case. Figure 1: A slice of the space of lightning self-attention mechanisms. *Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Identifiability of Deep Polynomial Neural NetworksKonstantin Usevich, Ricardo Augusto Borsoi, Clara Dérand, Marianne ClauselNeurIPS 2025 · 被引用 21 次
- Learning on a Razor's Edge: Identifiability and Singularity of Polynomial Neural NetworksVahid Shahverdi, Giovanni Luca Marchetti, Kathlén KohnICLR 2026 · 被引用 11 次
- Topology and geometry of the learning space of ReLU networks: connectivity and singularitiesMarco Nurisso, Pierrick Leroy, Giovanni Petri, Francesco VaccarinoICLR 2026 · 被引用 6 次
- Identifiable Equivariant Networks are Layerwise EquivariantVahid Shahverdi, Giovanni Luca Marchetti, Georg Bökman, Kathlén KohnICML 2026
它引用的顶会 Paper6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag, Kazuki Irie, Jürgen SchmidhuberICML 2021 · 被引用 394 次
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning DynamicsDaniel Kunin, Javier Sagastuy-Breña, Surya Ganguli, Daniel L. K. Yamins 等ICLR 2021 · 被引用 100 次
- Hidden Symmetries of ReLU NetworksJ. Elisenda Grigsby, Kathryn Lindsey, David RolnickICML 2023 · 被引用 35 次
- Symmetry Teleportation for Accelerated OptimizationBo Zhao, Nima Dehmamy, Robin Walters, Rose YuNeurIPS 2022 · 被引用 33 次
相关 Paper
- Attention Mechanism, Max-Affine Partition, and Universal ApproximationHude Liu, Jerry Yao-Chieh Hu, Zhao Song, Han LiuNeurIPS 2025 · 被引用 12 次
- Lipschitz normalization for self-attention layers with application to graph neural networksGeorge Dasoulas, Kevin Scaman, Aladin VirmauxICML 2021 · 被引用 55 次
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 被引用 172 次
- Spurious Valleys and Clustering Behavior of Neural NetworksSamuele PollaciICML 2023 · 被引用 1 次
- Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor ProgramsMana Sakai, Ryo Karakida, Masaaki ImaizumiNeurIPS 2025 · 被引用 5 次
