Weight decay induces low-rank attention layers
Seijin Kobayashi, Yassir Akram, Johannes von Oswald
Abstract
The effect of regularizers such as weight decay when training deep neural networks is not well understood. We study the influence of weight decay as well as -regularization when training neural network models in which parameter matrices interact multiplicatively. This combination is of particular interest as this parametrization is common in attention layers, the workhorse of transformers. Here, key-query, as well as value-projection parameter matrices, are multiplied directly with each other: and . We extend previous results and show on one hand that any local minimum of a -regularized loss of the form coincides with a minimum of the nuclear norm-regularized loss , and on the other hand that the 2 losses become identical exponentially quickly during training. We thus complement existing works linking -regularization with low-rank regularization, and in particular, explain why such regularization on the matrix product affects early stages of training. Based on these theoretical insights, we verify empirically that the key-query and value-projection matrix products within attention layers, when optimized with weight decay, as usually done in vision tasks and language modelling, indeed induce a significant reduction in the rank of and , even in fully online training. We find that, in accordance with existing work, inducing low rank in attention matrix products can damage language model performance, and observe advantages when decoupling weight decay in attention layers from the rest of the parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers14
- AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMsDi He, Songjun Tu, Ajay Jaiswal, Li Shen et al.NeurIPS 2025 · 14 citations
- Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling LawsFabrizio Boncoraglio, Vittorio Erba, Emanuele Troiani, Yizhou Xu et al.ICML 2026 · 5 citations
- Generalization Bounds for Rank-sparse Neural NetworksAntoine Ledent, Rodrigo Alves, Yunwen LeiNeurIPS 2025 · 4 citations
- Differentiable Sparsity via -Gating: Simple and Versatile Structured PenalizationChris Kolb, Laetitia Frost, Bernd Bischl, David RügamerNeurIPS 2025 · 4 citations
- Weight Decay Improves Language Model PlasticityTessa Han, Sebastian Bordt, Hanlin Zhang, Sham KakadeICML 2026 · 3 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Hyena Hierarchy: Towards Larger Convolutional Language ModelsMichael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu et al.ICML 2023 · 481 citations
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li et al.NeurIPS 2021 · 303 citations
Related papers
- Initialization and Regularization of Factorized Neural LayersMikhail Khodak, Neil A. Tenenholtz, Lester Mackey, Nicolò FusiICLR 2021 · 74 citations
- Understanding Decoupled and Early Weight DecayJohan Bjorck, Kilian Q. Weinberger, Carla P. GomesAAAI 2021 · 37 citations
- Why Do We Need Weight Decay in Modern Deep Learning?Francesco D'Angelo, Maksym Andriushchenko, Aditya Vardhan Varre, Nicolas FlammarionNeurIPS 2024 · 101 citations
- Understanding the Disharmony between Weight Normalization Family and Weight DecayXiang Li, Shuo Chen, Jian YangAAAI 2020 · 18 citations
- On the Role of Attention Masks and LayerNorm in TransformersXinyi Wu, Amir Ajorlou, Yifei Wang, Stefanie Jegelka et al.NeurIPS 2024 · 54 citations
