In defense of parameter sharing for model-compression
Aditya Desai, Anshumali Shrivastava
摘要
When considering a model architecture, there are several ways to reduce its memory footprint. Historically, popular approaches included selecting smaller architectures and creating sparse networks through pruning. More recently, randomized parameter-sharing (RPS) methods have gained traction for model compression at start of training. In this paper, we comprehensively assess the trade-off between memory and accuracy across RPS, pruning techniques, and building smaller models. Our findings demonstrate that RPS, which is both data and model-agnostic, consistently outperforms/matches smaller models and all moderately informed pruning strategies, such as MAG, SNIP, SYNFLOW, and GRASP, across the entire compression range. This advantage becomes particularly pronounced in higher compression scenarios. Notably, even when compared to highly informed pruning techniques like Lottery Ticket Rewinding (LTR), RPS exhibits superior performance in high compression settings. This points out inherent capacity advantage that RPS enjoys over sparse models. Theoretically, we establish RPS as a superior technique in terms of memory-efficient representation when compared to pruning for linear models. This paper argues in favor of paradigm shift towards RPS based models. During our rigorous evaluation of RPS, we identified issues in the stateof-the-art RPS technique ROAST, specifically regarding stability (ROAST's sensitivity to initialization hyperparameters, often leading to divergence) and Paretocontinuity (ROAST's inability to recover the accuracy of the original model at zero compression). We provably address both of these issues. We refer to the modified RPS, which incorporates our improvements, as STABLE-RPS. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Prune-then-Quantize or Quantize-then-Prune? Understanding the Impact of Compression Order in Joint Model CompressionMinjun Kim, Jaehyeon Choi, Hyunwoo Yang, Jongjin Kim 等ICLR 2026 · 被引用 5 次
- SS1: Accelerating Inference with Fast and Expressive Sketch Structured TransformAditya Desai, Kimia Saedi, Apoorv Walia, Jihyeong Lee 等NeurIPS 2024 · 被引用 1 次
它引用的顶会 Paper5
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- Prospect Pruning: Finding Trainable Weights at Initialization using Meta-GradientsMilad Alizadeh, Shyam A. Tailor, Luisa M. Zintgraf, Joost van Amersfoort 等ICLR 2022 · 被引用 50 次
- The trade-offs of model size in large recommendation models : 100GB to 10MB Criteo-tb DLRM modelAditya Desai, Anshumali ShrivastavaNeurIPS 2022 · 被引用 17 次
- Hardware-Aware Compression with Random Operation Access Specific Tile (ROAST) HashingAditya Desai, Keren Zhou, Anshumali ShrivastavaICML 2023 · 被引用 5 次
相关 Paper
- Robust Binary Models by Pruning Randomly-initialized NetworksChen Liu, Ziqi Zhao, Sabine Süsstrunk, Mathieu SalzmannNeurIPS 2022 · 被引用 7 次
- BackSlash: Rate Constrained Optimized Training of Large Language ModelsJun Wu, Jiangtao Wen, Yuxing HanICML 2025
- Why Random Pruning Is All We Need to Start SparseAdvait Harshal Gadhikar, Sohom Mukherjee, Rebekka BurkholzICML 2023 · 被引用 33 次
- Pruning Deep Neural Networks from a Sparsity PerspectiveEnmao Diao, Ganghua Wang, Jiawei Zhang, Yuhong Yang 等ICLR 2023 · 被引用 7 次
- How Well Do Sparse ImageNet Models Transfer?Eugenia Iofinova, Alexandra Peste, Mark Kurtz, Dan AlistarhCVPR 2022 · 被引用 20 次
