Differentiable Learning of Generalized Structured Matrices for Efficient Deep Neural Networks
Changwoo Lee, Hun-Seok Kim
摘要
This paper investigates efficient deep neural networks (DNNs) to replace dense unstructured weight matrices with structured ones that possess desired properties. The challenge arises because the optimal weight matrix structure in popular neural network models is obscure in most cases and may vary from layer to layer even in the same network. Prior structured matrices proposed for efficient DNNs were mostly hand-crafted without a generalized framework to systematically learn them. To address this issue, we propose a generalized and differentiable framework to learn efficient structures of weight matrices by gradient descent. We first define a new class of structured matrices that covers a wide range of structured matrices in the literature by adjusting the structural parameters. Then, the frequencydomain differentiable parameterization scheme based on the Gaussian-Dirichlet kernel is adopted to learn the structural parameters by proximal gradient descent. On the image and language tasks, our method learns efficient DNNs with structured matrices, achieving lower complexity and/or higher performance than prior approaches that employ low-rank, block-sparse, or block-low-rank matrices. 1. Is there a universal format that represents a wide range of structured matrices? 2. Can the structure of such matrices be learned efficiently, if it exists? Contributions. Tackling the above two questions, we introduce a generalized and differentiable structured matrix format. The main contributions of this work can be summarized as follows.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Compute Better Spent: Replacing Dense Layers with Structured MatricesShikai Qiu, Andres Potapczynski, Marc Anton Finzi, Micah Goldblum 等ICML 2024 · 被引用 26 次
- Searching for Efficient Linear Layers over a Continuous Space of Structured MatricesAndres Potapczynski, Shikai Qiu, Marc Finzi, Christopher Ferri 等NeurIPS 2024 · 被引用 11 次
- BLAST: Block-Level Adaptive Structured Matrices for Efficient Deep Neural Network InferenceChangwoo Lee, Soo Min Kwon, Qing Qu, Hun-Seok KimNeurIPS 2024 · 被引用 5 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 被引用 656 次
相关 Paper
- ASGO: Adaptive Structured Gradient OptimizationKang An, Yuxing Liu, Rui Pan, Yi Ren 等NeurIPS 2025 · 被引用 58 次
- Group and Shuffle: Efficient Structured Orthogonal ParametrizationMikhail Gorbunov, Nikolay Yudin, Vera Soboleva, Aibek Alanov 等NeurIPS 2024 · 被引用 11 次
- Tight Compression: Compressing CNN Model Tightly Through Unstructured Pruning and Simulated Annealing Based PermutationXizi Chen, Jingyang Zhu, Jingbo Jiang, Chi-Ying TsuiDAC 2020 · 被引用 30 次
- DARB: A Density-Adaptive Regular-Block Pruning for Deep Neural NetworksAo Ren, Tao Zhang, Yuhao Wang, Sheng Lin 等AAAI 2020 · 被引用 11 次
- The k-tied Normal Distribution: A Compact Parameterization of Gaussian Mean Field Posteriors in Bayesian Neural NetworksJakub Swiatkowski, Kevin Roth, Bastiaan S. Veeling, Linh Tran 等ICML 2020 · 被引用 52 次
