Attention layers provably solve single-location regression
Pierre Marion, Raphaël Berthier, Gérard Biau, Claire Boyer
摘要
Attention-based models, such as Transformer, excel across various tasks but lack a comprehensive theoretical understanding, especially regarding token-wise sparsity and internal linear representations. To address this gap, we introduce the single-location regression task, where only one token in a sequence determines the output, and its position is a latent random variable, retrievable via a linear projection of the input. To solve this task, we propose a dedicated predictor, which turns out to be a simplified version of a non-linear self-attention layer. We study its theoretical properties, by showing its asymptotic Bayes optimality and analyzing its training dynamics. In particular, despite the non-convex nature of the problem, the predictor effectively learns the underlying structure. This work highlights the capacity of attention mechanisms to handle sparse token information and internal linear structures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- The emergence of sparse attention: impact of data distribution and benefits of repetitionNicolas Zucchet, Francesco D'Angelo, Andrew Kyle Lampinen, Stephanie ChanNeurIPS 2025 · 被引用 28 次
- Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention NetworksLuca Arnaboldi, Bruno Loureiro, Ludovic Stephan, Florent Krzakala 等NeurIPS 2025 · 被引用 10 次
- High-Dimensional Analysis of Single-Layer Attention for Sparse-Token ClassificationNicholas Barnfield, Hugo Cui, Yue M. LuICLR 2026 · 被引用 8 次
- Statistical Advantage of Softmax Attention: Insights from Single-Location RegressionO. Duranthon, Pierre Marion, Claire Boyer, Bruno Loureiro 等ICLR 2026 · 被引用 7 次
- Bayes optimal learning of attention-indexed modelsFabrizio Boncoraglio, Emanuele Troiani, Vittorio Erba, Lenka ZdeborováNeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper17
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 被引用 324 次
相关 Paper
- How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman 等ICLR 2024 · 被引用 94 次
- Understanding Softmax Attention Layers: Exact Mean-Field Analysis on a Toy ProblemElvis DohmatobNeurIPS 2025 · 被引用 3 次
- One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-AttentionArvind V. Mahankali, Tatsunori Hashimoto, Tengyu MaICLR 2024 · 被引用 160 次
- How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear RegressionXingwu Chen, Lei Zhao, Difan ZouNeurIPS 2024 · 被引用 19 次
- How do Transformers Perform In-Context Autoregressive Learning ?Michael Eli Sander, Raja Giryes, Taiji Suzuki, Mathieu Blondel 等ICML 2024 · 被引用 21 次
