Focus and Dilution: The Multi-stage Learning Process of Attention
Zheng-An Chen, Pengxiao Lin, Zhi-Qin John Xu, Tao Luo
Abstract
Transformer-based models have achieved remarkable success across a wide range of domains, yet our understanding of their training dynamics remains limited. In this work, we identify a recurrent focus–dilution cycle in attention learning and provide a rigorous explanation in a one-layer Transformer setting for Markovian data via gradient-flow analysis. Using stage-wise linearization around critical points, we show that a single focus–dilution cycle can be decomposed into a sequence of distinct stages. First, embedding and projection rapidly condense to a rank-one structure, while attention parameters remain effectively frozen. Then, the attention parameters begin to increase, inducing a frequency-driven focus toward high-frequency tokens. As attention continues to evolve, it generates next-order perturbations in embeddings, leading to a mass-redistribution mechanism that progressively dilutes this focus. Finally, small asymmetries among low-frequency tokens lift a degenerate critical point, opening new embedding directions and initiating the next cycle. Experiments on synthetic Markovian data as well as WikiText and TinyStories corroborate the predicted stages and cyclical dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- Are Transformers universal approximators of sequence-to-sequence functions?Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J. Reddi et al.ICLR 2020 · 481 citations
- Birth of a Transformer: A Memory ViewpointAlberto Bietti, Vivien Cabannes, Diane Bouchacourt, Hervé Jégou et al.NeurIPS 2023 · 182 citations
- O(n) Connections are Expressive Enough: Universal Approximability of Sparse TransformersChulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat et al.NeurIPS 2020 · 111 citations
- How Do Transformers Learn Topic Structure: Towards a Mechanistic UnderstandingYuchen Li, Yuanzhi Li, Andrej RisteskiICML 2023 · 87 citations
- Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in TransformersSiyu Chen, Heejune Sheen, Tianhao Wang, Zhuoran YangNeurIPS 2024 · 48 citations
Related papers
- Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow AnalysisHongru Yang, Bhavya Kailkhura, Zhangyang Wang, Yingbin LiangNeurIPS 2024 · 14 citations
- JoMA: Demystifying Multilayer Transformers via Joint Dynamics of MLP and AttentionYuandong Tian, Yiping Wang, Zhenyu Zhang, Beidi Chen et al.ICLR 2024 · 49 citations
- Non-asymptotic Convergence of Training Transformers for Next-token PredictionRuiquan Huang, Yingbin Liang, Jing YangNeurIPS 2024 · 15 citations
- Incremental Learning of Sparse Attention Patterns in TransformersOğuz YükselICML 2026 · 2 citations
- On the Dynamics of Training Attention ModelsHaoye Lu, Yongyi Mao, Amiya NayakICLR 2021 · 9 citations
