Hardware-Aware Compression with Random Operation Access Specific Tile (ROAST) Hashing
Aditya Desai, Keren Zhou, Anshumali Shrivastava
Abstract
Advancements in deep learning are often associated with increasing model sizes. Training and deploying large models require sophisticated hardware and incur significantly higher costs. Thus, model compression is a widely explored approach to solving the problem. However, SOTA techniques fall short in one or more desirable aspects of compression -for instance, pruning does not reduce memory for training, quantization can only provide up to 32× compression, Hashed-Net is cache-inefficient, etc. This paper proposes a model-agnostic, cache-friendly, and hardwareaware model compression approach: Random Operation Access Specific Tile (ROAST) hashing. ROAST collapses the parameters by clubbing them through a lightweight mapping. While clubbing these parameters, ROAST utilizes cache hierarchies by aligning the memory access pattern with the parameter access pattern. ROAST is up to ∼25× faster to train and ∼50× faster to infer than the popular parameter sharing method HashedNet. Additionally, ROAST introduces global weight sharing, which is empirically and theoretically superior to local weight sharing in HashedNet, and can be of independent interest. With ROAST, we can efficiently train and deploy the model using a much smaller memory footprint (∼ 10 -100× lesser) in text and image classification tasks. ROAST-MM kernel implementation is open-source 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8584abfb-1384-4f89-afb3-a74da64c5e41Cited by top-tier papers3
- In defense of parameter sharing for model-compressionAditya Desai, Anshumali ShrivastavaICLR 2024 · 8 citations
- SS1: Accelerating Inference with Fast and Expressive Sketch Structured TransformAditya Desai, Kimia Saedi, Apoorv Walia, Jihyeong Lee et al.NeurIPS 2024 · 1 citation
- Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM AdaptationTianyi Zhang, Junda Su, Aditya Desai, Oscar Wu et al.ICML 2025
Builds on1
Related papers
- Structured Multi-Hashing for Model CompressionElad Eban, Yair Movshovitz-Attias, Hao Wu, Mark Sandler et al.CVPR 2020
- ALF: Autoencoder-based Low-rank Filter-sharing for Efficient Convolutional Neural NetworksAlexander Frickenstein, Manoj Rohit Vemparala, Nael Fasfous, Laura Hauenschild et al.DAC 2020 · 5 citations
- DRAGONN: Distributed Randomized Approximate Gradients of Neural NetworksZhuang Wang, Zhaozhuo Xu, Xinyu Crystal Wu, Anshumali Shrivastava et al.ICML 2022 · 10 citations
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly et al.AAAI 2021 · 79 citations
- Stitchable Neural NetworksZizheng Pan, Jianfei Cai, Bohan ZhuangCVPR 2023
