High-Dimensional Analysis of Single-Layer Attention for Sparse-Token Classification
Nicholas Barnfield, Hugo Cui, Yue M. Lu
摘要
When and how can an attention mechanism learn to selectively attend to informative tokens, thereby enabling detection of weak, rare, and sparsely located features? We address these questions theoretically in a sparse-token classification model in which positive samples embed a weak signal vector in a randomly chosen subset of tokens, whereas negative samples are pure noise. In the long-sequence limit, we show that a simple single-layer attention classifier can in principle achieve vanishing test error when the signal strength grows only logarithmically in the sequence length , whereas linear classifiers require scaling. Moving from representational power to learnability, we study training at finite in a high-dimensional regime, where sample size and embedding dimension grow proportionally. We prove that just two gradient updates suffice for the query weight vector of the attention classifier to acquire a nontrivial alignment with the hidden signal, inducing an attention map that selectively amplifies informative tokens. We further derive an exact asymptotic expression for the test error and training loss of the trained attention-based classifier, and quantify its capacity -- the largest dataset size that is typically perfectly separable -- thereby explaining the advantage of adaptive token selection over nonadaptive linear baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Statistical Advantage of Softmax Attention: Insights from Single-Location RegressionO. Duranthon, Pierre Marion, Claire Boyer, Bruno Loureiro 等ICLR 2026 · 被引用 7 次
- Provably Learning Attention with QueriesSatwik Bhattamishra, Kulin Shah, Michael Hahn, Varun KanadeICML 2026
- Asymptotic Theory of Iterated Empirical Risk Minimization, with Applications to Active LearningHugo Cui, Yue LuICML 2026
它引用的顶会 Paper31
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento 等ICML 2023 · 被引用 729 次
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 被引用 324 次
- Birth of a Transformer: A Memory ViewpointAlberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou 等NeurIPS 2023 · 被引用 182 次
相关 Paper
- Benign Overfitting in Token Selection of Attention MechanismKeitaro Sakamoto, Issei SatoICML 2025
- Attention layers provably solve single-location regressionPierre Marion, Raphaël Berthier, Gérard Biau, Claire BoyerICLR 2025
- Limitations of Normalization in AttentionTimur Mudarisov, Mikhail Burtsev, Tatiana Petrova, Radu StateNeurIPS 2025
- When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical PerspectiveAlireza Mousavi-Hosseini, Clayton Sanford, Denny Wu, Murat A. ErdogduNeurIPS 2025 · 被引用 6 次
- Inductive Biases and Variable Creation in Self-Attention MechanismsBenjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Cyril ZhangICML 2022 · 被引用 154 次
