Structured Sparsification of Gated Recurrent Neural Networks
Ekaterina Lobacheva, Nadezhda Chirkova, Alexander Markovich, Dmitry P. Vetrov
Abstract
One of the most popular approaches for neural network compression is sparsification — learning sparse weight matrices. In structured sparsification, weights are set to zero by groups corresponding to structure units, e. g. neurons. We further develop the structured sparsification approach for the gated recurrent neural networks, e. g. Long Short-Term Memory (LSTM). Specifically, in addition to the sparsification of individual weights and neurons, we propose sparsifying the preactivations of gates. This makes some gates constant and simplifies an LSTM structure. We test our approach on the text classification and language modeling tasks. Our method improves the neuron-wise compression of the model in most of the tasks. We also observe that the resulting structure of gate sparsity depends on the task and connect the learned structures to the specifics of the particular tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Selfish Sparse RNN TrainingShiwei Liu, Decebal Constantin Mocanu, Yulong Pei, Mykola PechenizkiyICML 2021 · 43 citations
- Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient TrainingAnup Sarma, Sonali Singh, Huaipan Jiang, Rui Zhang et al.NeurIPS 2021 · 1 citation
- One-Shot Pruning of Recurrent Neural Networks by Jacobian Spectrum EvaluationMatthew Shunshi Zhang, Bradly C. StadieICLR 2020 · 34 citations
- SUBP: Soft Uniform Block Pruning for 1×N Sparse CNNs Multithreading AccelerationJingyang Xiang, Siqi Li, Jun Chen, Guang Dai et al.NeurIPS 2023 · 2 citations
- Differentiable Sparsity via -Gating: Simple and Versatile Structured PenalizationChris Kolb, Laetitia Frost, Bernd Bischl, David RügamerNeurIPS 2025 · 4 citations
