Geometric Signatures of Compositionality Across a Language Model's Lifetime
Jin Hwa Lee, Thomas Jiralerspong, Lei Yu, Yoshua Bengio, Emily Cheng
Abstract
By virtue of linguistic compositionality, few syntactic rules and a finite lexicon can generate an unbounded number of sentences. That is, language, though seemingly high-dimensional, can be explained using relatively few degrees of freedom. An open question is whether contemporary language models (LMs) reflect the intrinsic simplicity of language that is enabled by compositionality. We take a geometric view of this problem by relating the degree of compositionality in a dataset to the intrinsic dimension (I d ) of its representations under an LM, a measure of feature complexity. We find not only that the degree of dataset compositionality is reflected in representations' I d , but that the relationship between compositionality and geometric complexity arises due to learned linguistic features over training. Finally, our analyses reveal a striking contrast between nonlinear I d and linear dimensionality, showing they respectively encode semantic and superficial aspects of linguistic composition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 18f0912b-4836-40f3-aacf-a66b6f34caa2Cited by top-tier papers4
- From Directions to Regions: Decomposing Activations in Language Models via Local GeometryOr Shafran, Shaked Ronen, Omri Fahn, Shauli Ravfogel et al.ICML 2026 · 7 citations
- Less is More: Local Intrinsic Dimensions of Contextual Language ModelsBenjamin Matthias Ruppik, Julius von Rohrscheidt, Carel van Niekerk, Michael Heck et al.NeurIPS 2025 · 1 citation
- From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based ModelsLuca Dini, Lucia Domenichelli, Dominique Brunato, Felice Dell'OrlettaACL 2025
- Abstraction Induces the Brain Alignment of Language and Speech ModelsEmily Cheng, Aditya Vaidya, Richard AntonelloICML 2026
Builds on19
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- The Intrinsic Dimension of Images and Its Impact on LearningPhillip Pope, Chen Zhu, Ahmed Abdelkader, Micah Goldblum et al.ICLR 2021 · 381 citations
- Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMsAngelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt et al.ICLR 2024 · 119 citations
- The Transient Nature of Emergent In-Context Learning in TransformersAaditya K. Singh, Stephanie C. Y. Chan, Ted Moskovitz, Erin Grant et al.NeurIPS 2023 · 92 citations
Related papers
- Bridging Information-Theoretic and Geometric Compression in Language ModelsEmily Cheng, Corentin Kervadec, Marco BaroniEMNLP 2023 · 5 citations
- Emergence of a High-Dimensional Abstraction Phase in Language TransformersEmily Cheng, Diego Doimo, Corentin Kervadec, Iuri Macocco et al.ICLR 2025 · 1 citation
- Are representations built from the ground up? An empirical examination of local composition in language modelsEmmy Liu, Graham NeubigEMNLP 2022 · 5 citations
- Geometry of Decision Making in Language ModelsAbhinav Joshi, Divyanshu Bhatt, Ashutosh ModiNeurIPS 2025 · 12 citations
- Correlation Dimension of Autoregressive Large Language ModelsXin Du, Kumiko Tanaka-IshiiNeurIPS 2025 · 2 citations
