In defense of parameter sharing for model-compression
Aditya Desai, Anshumali Shrivastava
Abstract
When considering a model architecture, there are several ways to reduce its memory footprint. Historically, popular approaches included selecting smaller architectures and creating sparse networks through pruning. More recently, randomized parameter-sharing (RPS) methods have gained traction for model compression at start of training. In this paper, we comprehensively assess the trade-off between memory and accuracy across RPS, pruning techniques, and building smaller models. Our findings demonstrate that RPS, which is both data and model-agnostic, consistently outperforms/matches smaller models and all moderately informed pruning strategies, such as MAG, SNIP, SYNFLOW, and GRASP, across the entire compression range. This advantage becomes particularly pronounced in higher compression scenarios. Notably, even when compared to highly informed pruning techniques like Lottery Ticket Rewinding (LTR), RPS exhibits superior performance in high compression settings. This points out inherent capacity advantage that RPS enjoys over sparse models. Theoretically, we establish RPS as a superior technique in terms of memory-efficient representation when compared to pruning for linear models. This paper argues in favor of paradigm shift towards RPS based models. During our rigorous evaluation of RPS, we identified issues in the stateof-the-art RPS technique ROAST, specifically regarding stability (ROAST's sensitivity to initialization hyperparameters, often leading to divergence) and Paretocontinuity (ROAST's inability to recover the accuracy of the original model at zero compression). We provably address both of these issues. We refer to the modified RPS, which incorporates our improvements, as STABLE-RPS. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d5ba064-e789-4570-ba13-df098b9b7319Cited by top-tier papers2
- Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model CompressionMinjun Kim, Jaehyeon Choi, Hyunwoo Yang, Jongjin Kim et al.ICLR 2026 · 5 citations
- SS1: Accelerating Inference with Fast and Expressive Sketch Structured TransformAditya Desai, Kimia Saedi, Apoorv Walia, Jihyeong Lee et al.NeurIPS 2024 · 1 citation
Builds on5
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Prospect Pruning: Finding Trainable Weights at Initialization using Meta-GradientsMilad Alizadeh, Shyam A. Tailor, Luisa M. Zintgraf, Joost van Amersfoort et al.ICLR 2022 · 50 citations
- The trade-offs of model size in large recommendation models : 100GB to 10MB Criteo-tb DLRM modelAditya Desai, Anshumali ShrivastavaNeurIPS 2022 · 17 citations
- Hardware-Aware Compression with Random Operation Access Specific Tile (ROAST) HashingAditya Desai, Keren Zhou, Anshumali ShrivastavaICML 2023 · 5 citations
Related papers
- Robust Binary Models by Pruning Randomly-initialized NetworksChen Liu, Ziqi Zhao, Sabine Süsstrunk, Mathieu SalzmannNeurIPS 2022 · 7 citations
- BackSlash: Rate Constrained Optimized Training of Large Language ModelsJun Wu, Jiangtao Wen, Yuxing HanICML 2025
- Why Random Pruning Is All We Need to Start SparseAdvait Harshal Gadhikar, Sohom Mukherjee, Rebekka BurkholzICML 2023 · 33 citations
- Pruning Deep Neural Networks from a Sparsity PerspectiveEnmao Diao, Ganghua Wang, Jiawei Zhang, Yuhong Yang et al.ICLR 2023 · 7 citations
- How Well Do Sparse ImageNet Models Transfer?Eugenia Iofinova, Alexandra Peste, Mark Kurtz, Dan AlistarhCVPR 2022 · 20 citations
