One-Shot Pruning of Recurrent Neural Networks by Jacobian Spectrum Evaluation
Matthew Shunshi Zhang, Bradly C. Stadie
摘要
Recent advances in the sparse neural network literature have made it possible to prune many large feed forward and convolutional networks with only a small quantity of data. Yet, these same techniques often falter when applied to the problem of recovering sparse recurrent networks. These failures are quantitative: when pruned with recent techniques, RNNs typically obtain worse performance than they do under a simple random pruning scheme. The failures are also qualitative: the distribution of active weights in a pruned LSTM or GRU network tend to be concentrated in specific neurons and gates, and not well dispersed across the entire architecture. We seek to rectify both the quantitative and qualitative issues with recurrent network pruning by introducing a new recurrent pruning objective derived from the spectrum of the recurrent Jacobian. Our objective is data efficient (requiring only 64 data points to prune the network), easy to implement, and produces 95 % sparse GRUs that significantly improve on existing baselines. We evaluate on sequential MNIST, Billion Words, and Wikitext.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly 等AAAI 2021 · 被引用 79 次
- KD-Zero: Evolving Knowledge Distiller for Any Teacher-Student PairsLujun Li, Peijie Dong, Anggeng Li, Zimian Wei 等NeurIPS 2023 · 被引用 49 次
相关 Paper
- Selfish Sparse RNN TrainingShiwei Liu, Decebal Constantin Mocanu, Yulong Pei, Mykola PechenizkiyICML 2021 · 被引用 43 次
- Structured Sparsification of Gated Recurrent Neural NetworksEkaterina Lobacheva, Nadezhda Chirkova, Alexander Markovich, Dmitry P. VetrovAAAI 2020 · 被引用 3 次
- Sparse Spiking Neural Network: Exploiting Heterogeneity in Timescales for Pruning Recurrent SNNBiswadeep Chakraborty, Beomseok Kang, Harshit Kumar, Saibal MukhopadhyayICLR 2024 · 被引用 18 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Sparse Deep Learning for Time Series Data: Theory and ApplicationsMingxuan Zhang, Yan Sun, Faming LiangNeurIPS 2023 · 被引用 10 次
