Mechanism and Emergence of Stacked Attention Heads in Multi-Layer Transformers
Tiberiu Musat
Abstract
In this paper, I introduce the retrieval problem, a simple yet common reasoning task that can be solved only by transformers with a minimum number of layers, which grows logarithmically with the input size. I empirically show that large language models can solve the task under different prompting formulations without any fine-tuning. To understand how transformers solve the retrieval problem, I train several transformers on a minimal formulation. Successful learning occurs only under the presence of an implicit curriculum. I uncover the learned mechanisms by studying the attention maps in the trained transformers. I also study the training process, uncovering that attention heads always emerge in a specific sequence guided by the implicit curriculum.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- Delta Attention: Fast and Accurate Sparse Attention Inference by Delta CorrectionJeffrey Willette, Heejun Lee, Sung Ju HwangNeurIPS 2025 · 9 citations
- Stability Implies Redundancy: Delta Attention Selective Halting for Efficient Long-Context PrefillingYujie Chen, Tailai Chen, Yifeng Gao, Zoe Wanying He et al.ACL 2026 · 1 citation
Related papers
- Language models can learn implicit multi-hop reasoning, but only if they have lots of training dataYuekun Yao, Yupei Du, Dawei Zhu, Michael Hahn et al.EMNLP 2025
- How do Transformers Learn Implicit Reasoning?Jiaran Ye, Zijun Yao, Zhidian Huang, Liangming Pan et al.NeurIPS 2025 · 17 citations
- Retrieval Head Mechanistically Explains Long-Context FactualityWenhao Wu, Yizhong Wang, Guangxuan Xiao, Hao Peng et al.ICLR 2025
- Why Prompt Design Matters and Works: A Complexity Analysis of Prompt Search Space in LLMsXiang Zhang, Juntai Cao, Chenyu You, Dujian DingACL 2025 · 21 citations
- Exploring Length Generalization in Large Language ModelsCem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz et al.NeurIPS 2022 · 267 citations
