Initialization and Regularization of Factorized Neural Layers
Mikhail Khodak, Neil A. Tenenholtz, Lester Mackey, Nicolò Fusi
Abstract
Factorized layers--operations parameterized by products of two or more matrices--occur in a variety of deep learning contexts, including compressed model training, certain types of knowledge distillation, and multi-head self-attention architectures. We study how to initialize and regularize deep nets containing such layers, examining two simple, understudied schemes, spectral initialization and Frobenius decay, for improving their performance. The guiding insight is to design optimization routines for these networks that are as close as possible to that of their well-tuned, non-decomposed counterparts; we back this intuition with an analysis of how the initialization and regularization schemes impact training with gradient descent, drawing on modern attempts to understand the interplay of weight-decay and batch-normalization. Empirically, we highlight the benefits of spectral initialization and Frobenius decay across a variety of settings. In model compression, we show that they enable low-rank methods to significantly outperform both unstructured sparsity and tensor methods on the task of training low-memory residual networks; analogs of the schemes also improve the performance of tensor decomposition techniques. For knowledge distillation, Frobenius decay enables a simple, overcomplete baseline that yields a compact model from over-parameterized training without requiring retraining with or pruning a teacher network. Finally, we show how both schemes applied to multi-head attention lead to improved performance on both translation and unsupervised pre-training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb8a939b-e74f-459d-84f4-e1f1b5d04f6fCited by top-tier papers30
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- DRONE: Data-aware Low-rank Compression for Large NLP ModelsPatrick H. Chen, Hsiang-Fu Yu, Inderjit S. Dhillon, Cho-Jui HsiehNeurIPS 2021 · 109 citations
- Low-rank lottery tickets: finding efficient low-rank neural networks via matrix differential equationsSteffen Schotthöfer, Emanuele Zangrando, Jonas Kusch, Gianluca Ceruti et al.NeurIPS 2022 · 66 citations
- Hypernetwork-based Meta-Learning for Low-Rank Physics-Informed Neural NetworksWoojin Cho, Kookjin Lee, Donsub Rim, Noseong ParkNeurIPS 2023 · 62 citations
- SLTrain: a sparse plus low rank approach for parameter and memory efficient pretrainingAndi Han, Jiaxiang Li, Wei Huang, Mingyi Hong et al.NeurIPS 2024 · 54 citations
Builds on9
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution BlocksXiaohan Ding, Yuchen Guo, Guiguang Ding, Jungong HanICCV 2019 · 845 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
Related papers
- Shapeshifter: a Parameter-efficient Transformer using Factorized Reshaped MatricesAliakbar Panahi, Seyran Saeedi, Tom ArodzNeurIPS 2021 · 19 citations
- Weight decay induces low-rank attention layersSeijin Kobayashi, Yassir Akram, Johannes von OswaldNeurIPS 2024 · 41 citations
- Geometry-aware training of factorized layers in tensor Tucker formatEmanuele Zangrando, Steffen Schotthöfer, Gianluca Ceruti, Jonas Kusch et al.NeurIPS 2024 · 20 citations
- Batch normalization provably avoids ranks collapse for randomly initialised deep networksHadi Daneshmand, Jonas Moritz Kohler, Francis R. Bach, Thomas Hofmann et al.NeurIPS 2020 · 73 citations
- Deep Weight Factorization: Sparse Learning Through the Lens of Artificial SymmetriesChris Kolb, Tobias Weber, Bernd Bischl, David RügamerICLR 2025
