How do Transformers Perform In-Context Autoregressive Learning ?
Michael Eli Sander, Raja Giryes, Taiji Suzuki, Mathieu Blondel, Gabriel Peyré
摘要
Transformers have achieved state-of-the-art performance in language modeling tasks. However, the reasons behind their tremendous success are still unclear. In this paper, towards a better understanding, we train a Transformer model on a simple next token prediction task, where sequences are generated as a first-order autoregressive process . We show how a trained Transformer predicts the next token by first learning in-context, then applying a prediction mapping. We call the resulting procedure in-context autoregressive learning. More precisely, focusing on commuting orthogonal matrices , we first show that a trained one-layer linear Transformer implements one step of gradient descent for the minimization of an inner objective function, when considering augmented tokens. When the tokens are not augmented, we characterize the global minima of a one-layer diagonal linear multi-head Transformer. Importantly, we exhibit orthogonality between heads and show that positional encoding captures trigonometric relations in the data. On the experimental side, we consider the general case of non-commuting orthogonal matrices and generalize our theoretical findings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Pretrained Transformer Efficiently Learns Low-Dimensional Target Functions In-ContextKazusato Oko, Yujin Song, Taiji Suzuki, Denny WuNeurIPS 2024 · 被引用 34 次
- On Mesa-Optimization in Autoregressively Trained Transformers: Emergence and CapabilityChenyu Zheng, Wei Huang, Rongzhen Wang, Guoqiang Wu 等NeurIPS 2024 · 被引用 10 次
- Transformers are Universal In-context LearnersTakashi Furuya, Maarten V. de Hoop, Gabriel PeyréICLR 2025 · 被引用 2 次
- Transformers Learn Latent Mixture Models In-Context via Mirror DescentFrancesco D'Angelo, Nicolas FlammarionICLR 2026 · 被引用 2 次
- Token Sample Complexity of AttentionLéa Bohbot, Cyril Letrouit, Gabriel Peyré, François-Xavier VialardICML 2026 · 被引用 1 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
相关 Paper
- How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear RegressionXingwu Chen, Lei Zhao, Difan ZouNeurIPS 2024 · 被引用 19 次
- Towards Understanding the Universality of Transformers for Next-Token PredictionMichael Eli Sander, Gabriel PeyréICLR 2025
- How Transformers Learn Causal Structure with Gradient DescentEshaan Nichani, Alex Damian, Jason D. LeeICML 2024 · 被引用 117 次
- Local to Global: Learning Dynamics and Effect of Initialization for TransformersAshok Vardhan Makkuva, Marco Bondaschi, Adway Girish, Alliot Nagle 等NeurIPS 2024 · 被引用 16 次
- How Transformers Represent Hierarchies: A Local-to-Global MechanismZhiling Zhou, Tianhao Wang, Zhuoran YangICML 2026
