Transformers Represent Belief State Geometry in their Residual Stream
Adam S. Shai, Lucas Teixeira, Alexander Gietelink Oldenziel, Sarah Marzen, Paul M. Riechers
摘要
What computational structure are we building into large language models when we train them on next-token prediction? Here, we present evidence that this structure is given by the meta-dynamics of belief updating over hidden states of the data-generating process. Leveraging the theory of optimal prediction, we anticipate and then find that belief states are linearly represented in the residual stream of transformers, even in cases where the predicted belief state geometry has highly nontrivial fractal structure. We investigate cases where the belief state geometry is represented in the final residual stream or distributed across the residual streams of multiple layers, providing a framework to explain these observations. Furthermore we demonstrate that the inferred belief states contain information about the entire future, beyond the local next-token prediction that the transformers are explicitly trained on. Our work provides a general framework connecting the structure of training data to the geometric structure of activations inside transformers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Symmetries in language statistics shape the geometry of model representationsDhruva Karkada, Daniel Korchinski, Andres Nava, Matthieu Wyart 等ICML 2026 · 被引用 15 次
- Not All Language Model Features Are One-Dimensionally LinearJoshua Engels, Eric J. Michaud, Isaac Liao, Wes Gurnee 等ICLR 2025 · 被引用 3 次
- Spectral Filters, Dark Signals, and Attention SinksNicola CanceddaACL 2024 · 被引用 3 次
- Transformers learn factored representationsAdam Shai, Loren Amdahl-Culleton, Casper Christensen, Henry R Bigelow 等ICML 2026 · 被引用 2 次
- Transformers Learn Latent Mixture Models In-Context via Mirror DescentFrancesco D'Angelo, Nicolas FlammarionICLR 2026 · 被引用 2 次
它引用的顶会 Paper4
- Better & Faster Large Language Models via Multi-token PredictionFabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz 等ICML 2024 · 被引用 286 次
- The Pitfalls of Next-Token PredictionGregor Bachmann, Vaishnavh NagarajanICML 2024 · 被引用 163 次
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic TaskKenneth Li, Aspen K. Hopkins, David Bau, Fernanda B. Viégas 等ICLR 2023 · 被引用 60 次
- Transformers Learn Shortcuts to AutomataBingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy 等ICLR 2023 · 被引用 11 次
相关 Paper
- Constrained Belief Updates Explain Geometric Structures in Transformer RepresentationsMateusz Piotrowski, Paul M. Riechers, Daniel Filan, Adam S. ShaiICML 2025
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
- Provable Long-Range Benefits of Next-Token PredictionXinyuan Cao, Santosh S. VempalaSTOC 2026
- Residual Stream Analysis with Multi-Layer SAEsTim Lawson, Lucy Farnik, Conor J. Houghton, Laurence AitchisonICLR 2025
- Lines of Thought in Large Language ModelsRaphaël Sarfati, Toni J. B. Liu, Nicolas Boullé, Christopher J. EarlsICLR 2025
