Transformers learn factored representations
Adam Shai, Loren Amdahl-Culleton, Casper Christensen, Henry R Bigelow, Fernando Rosas, Alexander Boyd, Eric Alt, Kyle Ray, Paul Riechers
摘要
Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize two representational hypotheses: (1) a representation in the product space of all factors, whose dimension grows exponentially with the number of parts, or (2) a factored representation in orthogonal subspaces, whose dimension grows linearly. The factored representation is lossless when factors are conditionally independent, but sacrifices predictive fidelity otherwise, creating a tradeoff between dimensional efficiency and accuracy. We derive precise predictions about the geometric structure of activations for each, including the number of subspaces, their dimensionality, and the arrangement of context embeddings within them. We test between these hypotheses on transformers trained on synthetic processes with known latent structure. Models learn factored representations when factors are conditionally independent, and continue to favor them early in training even when noise or hidden dependencies undermine conditional independence, reflecting an inductive bias toward factoring at the cost of fidelity. This provides a principled explanation for why transformers decompose the world into parts, and suggests that interpretable low dimensional structure may persist even in models trained on complex data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Transformers Represent Belief State Geometry in their Residual StreamAdam S. Shai, Lucas Teixeira, Alexander Gietelink Oldenziel, Sarah Marzen 等NeurIPS 2024 · 被引用 83 次
- Dense SAE Latents Are Features, Not BugsXiaoqing Sun, Alessandro Stolfo, Joshua Engels, Ben Wu 等NeurIPS 2025 · 被引用 19 次
- Disentangling Representations of Text by Masking TransformersXiongyi Zhang, Jan-Willem van de Meent, Byron C. WallaceEMNLP 2021 · 被引用 12 次
- Transformers Learn Shortcuts to AutomataBingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy 等ICLR 2023 · 被引用 11 次
- Not All Language Model Features Are One-Dimensionally LinearJoshua Engels, Eric J. Michaud, Isaac Liao, Wes Gurnee 等ICLR 2025 · 被引用 3 次
相关 Paper
- Constrained Belief Updates Explain Geometric Structures in Transformer RepresentationsMateusz Piotrowski, Paul M. Riechers, Daniel Filan, Adam S. ShaiICML 2025
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
- Understanding the Emergence of Seemingly Useless Features in Next-Token PredictorsMark Rofin, Jalal Naghiyev, Michael HahnICLR 2026
- Emergence of Linear Truth Encodings in Language ModelsShauli Ravfogel, Gilad Yehudai, Tal Linzen, Joan Bruna 等NeurIPS 2025 · 被引用 12 次
- On the Origins of Linear Representations in Large Language ModelsYibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam 等ICML 2024 · 被引用 68 次
