Low-Rank Bottleneck in Multi-head Attention Models
Srinadh Bhojanapalli, Chulhee Yun, Ankit Singh Rawat, Sashank J. Reddi, Sanjiv Kumar
摘要
Attention based Transformer architecture has enabled significant advances in the field of natural language processing. In addition to new pre-training techniques, recent improvements crucially rely on working with a relatively larger embedding dimension for tokens. Unfortunately, this leads to models that are prohibitively large to be employed in the downstream tasks. In this paper we identify one of the important factors contributing to the large embedding size requirement. In particular, our analysis highlights that the scaling between the number of heads and the size of each head in the current architecture gives rise to a low-rank bottleneck in attention heads, causing this limitation. We further validate this in our experiments. As a solution we propose to set the head size of an attention unit to input sequence length, and independent of the number of heads, resulting in multi-head attention layers with provably more expressive power. We empirically show that this allows us to train models with a relatively smaller embedding dimension and with better performance scaling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper50
- YOLOv12: Attention-Centric Real-Time Object DetectorsYunjie Tian, Qixiang Ye, David S. DoermannNeurIPS 2025 · 被引用 2,652 次
- Pure Transformers are Powerful Graph LearnersJinwoo Kim, Dat Nguyen, Seonwoo Min, Sungjun Cho 等NeurIPS 2022 · 被引用 311 次
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 被引用 214 次
- ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accelerator Co-DesignHaoran You, Zhanyi Sun, Huihong Shi, Zhongzhi Yu 等HPCA 2023 · 被引用 124 次
- O(n) Connections are Expressive Enough: Universal Approximability of Sparse TransformersChulhee Yun, Yin-Wen Chang, Srinadh Bhojanapalli, Ankit Singh Rawat 等NeurIPS 2020 · 被引用 111 次
它引用的顶会 Paper1
相关 Paper
- Quality over Quantity in Attention Layers: When Adding More Heads HurtsNoah Amsel, Gilad Yehudai, Joan BrunaICLR 2025
- Which transformer architecture fits my data? A vocabulary bottleneck in self-attentionNoam Wies, Yoav Levine, Daniel Jannai, Amnon ShashuaICML 2021 · 被引用 22 次
- Two Heads are Better than One: Simulating Large Transformers with Small OnesHantao Yu, Josh AlmanNeurIPS 2025 · 被引用 1 次
- Representational Strengths and Limitations of TransformersClayton Sanford, Daniel J. Hsu, Matus TelgarskyNeurIPS 2023 · 被引用 162 次
- Scalable Efficient Training of Large Language Models with Low-dimensional Projected AttentionXingtai Lv, Ning Ding, Kaiyan Zhang, Ermo Hua 等EMNLP 2024 · 被引用 2 次
