Symmetries in language statistics shape the geometry of model representations
Dhruva Karkada, Daniel Korchinski, Andres Nava, Matthieu Wyart, Yasaman Bahri
Abstract
The internal representations learned by language models consistently exhibit striking geometric structure: calendar months organize into a circle, historical years form a smooth one-dimensional manifold, and cities' latitudes and longitudes can be decoded using a linear probe. To explain this neural code, we first show that language statistics exhibit translation symmetry (for example, the frequency with which any two months co-occur in text depends only on the time interval between them). We prove that this symmetry governs these geometric structures in high-dimensional word embedding models, and we analytically derive the manifold geometry of word representations. These predictions empirically match large text embedding models and large language models. Moreover, the representational geometry persists at moderate embedding dimension even when the relevant statistics are perturbed (e.g., by removing all sentences in which two months co-occur). We prove that this robustness emerges naturally when the co-occurrence statistics are controlled by an underlying latent variable. Our results indicate that these representational manifolds originate in the statistical symmetries of natural language.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d223ad70-9995-4467-bbbf-6bb899ed1a29Builds on18
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 303 citations
- Birth of a Transformer: A Memory ViewpointAlberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou et al.NeurIPS 2023 · 182 citations
- The Evolution of Statistical Induction Heads: In-Context Learning Markov ChainsEzra Edelman, Nikolaos Tsilivis, Benjamin L. Edelman, Eran Malach et al.NeurIPS 2024 · 140 citations
- How Transformers Learn Causal Structure with Gradient DescentEshaan Nichani, Alex Damian, Jason D. LeeICML 2024 · 117 citations
Related papers
- Low-dimensional Structure in the Space of Language Representations is Reflected in Brain ResponsesRichard J. Antonello, Javier S. Turek, Vy Ai Vo, Alexander HuthNeurIPS 2021 · 60 citations
- Isotropy in the Contextual Embedding Space: Clusters and ManifoldsXingyu Cai, Jiaji Huang, Yuchen Bian, Kenneth ChurchICLR 2021 · 50 citations
- Connecting Neural Models Latent Geometries with Relative Geodesic RepresentationsHanlin Yu, Berfin Inal, Georgios Arvanitidis, Søren Hauberg et al.NeurIPS 2025 · 5 citations
- Geometric Signatures of Compositionality Across a Language Model's LifetimeJin Hwa Lee, Thomas Jiralerspong, Lei Yu, Yoshua Bengio et al.ACL 2025
- Lines of Thought in Large Language ModelsRaphaël Sarfati, Toni J. B. Liu, Nicolas Boullé, Christopher J. EarlsICLR 2025
