Large language models implicitly learn to straighten neural sentence trajectories to construct a predictive representation of natural language
Eghbal A. Hosseini, Evelina Fedorenko
Abstract
Predicting upcoming events is critical to our ability to effectively interact with our environment and conspecifics. In natural language processing, transformer models, which are trained on next-word prediction, appear to construct a general-purpose representation of language that can support diverse downstream tasks. However, we still lack an understanding of how a predictive objective shapes such representations. Inspired by recent work in vision neuroscience Hénaff et al. (2019) , here we test a hypothesis about predictive representations of autoregressive transformer models. In particular, we test whether the neural trajectory of a sequence of words in a sentence becomes progressively more straight as it passes through the layers of the network. The key insight behind this hypothesis is that straighter trajectories should facilitate prediction via linear extrapolation. We quantify straightness using a 1dimensional curvature metric, and present four findings in support of the trajectory straightening hypothesis: i) In trained models, the curvature progressively decreases from the first to the middle layers of the network. ii) Models that perform better on the next-word prediction objective, including larger models and models trained on larger datasets, exhibit greater decreases in curvature, suggesting that this improved ability to straighten sentence neural trajectories may be the underlying driver of better language modeling performance. iii) Given the same linguistic context, the sequences that are generated by the model have lower curvature than the ground truth (the actual continuations observed in a language corpus), suggesting that the model favors straighter trajectories for making predictions. iv) A consistent relationship holds between the average curvature and the average surprisal of sentences in the middle layers of models, such that sentences with straighter neural trajectories also have lower surprisal. Importantly, untrained models don't exhibit these behaviors. In tandem, these results support the trajectory straightening hypothesis and provide a possible mechanism for how the geometry of the internal representations of autoregressive models supports next word prediction. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b5f6af0c-64ae-4906-8aab-401bf439208aCited by top-tier papers14
- LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut ModulationAhmadreza Jeddi, Marco Ciccone, Babak TaatiICLR 2026 · 54 citations
- AI-Generated Video Detection via Perceptual StraighteningChristian Internò, Robert Geirhos, Markus Olhofer, Sunny Liu et al.NeurIPS 2025 · 44 citations
- Priors in time: Missing inductive biases for language model interpretabilityEkdeep Singh Lubana, Can Rager, Sai Sumedh R. Hindupur, Valérie Costa et al.ICLR 2026 · 19 citations
- Temporal Straightening for Latent PlanningYing Wang, Oumayma Bounou, Gaoyue Zhou, Randall Balestriero et al.ICML 2026 · 19 citations
- Tracing the Traces: Latent Temporal Signals for Efficient and Accurate ReasoningMartina G. Vilas, Safoora Yousefi, Besmira Nushi, Eric Horvitz et al.ICLR 2026 · 12 citations
Builds on3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The geometry of hidden representations of large transformer modelsLucrezia Valeriani, Diego Doimo, Francesca Cuturello, Alessandro Laio et al.NeurIPS 2023 · 148 citations
- Emergence of Separable Manifolds in Deep Language RepresentationsJonathan Mamou, Hang Le, Miguel Del Rio, Cory Stephenson et al.ICML 2020 · 52 citations
Related papers
- Representational Curvature Modulates Behavioral Uncertainty in Large Language ModelsJack King, Evelina Fedorenko, Eghbal HosseiniICML 2026
- How do Transformers Perform In-Context Autoregressive Learning ?Michael Eli Sander, Raja Giryes, Taiji Suzuki, Mathieu Blondel et al.ICML 2024 · 21 citations
- A Mathematical Exploration of Why Language Models Help Solve Downstream TasksNikunj Saunshi, Sadhika Malladi, Sanjeev AroraICLR 2021 · 93 citations
- Exploring perceptual straightness in learned visual representationsAnne Harrington, Vasha DuTell, Ayush Tewari, Mark Hamilton et al.ICLR 2023
- Provable Long-Range Benefits of Next-Token PredictionXinyuan Cao, Santosh S. VempalaSTOC 2026
