Hardware-Aware Compression with Random Operation Access Specific Tile (ROAST) Hashing
Aditya Desai, Keren Zhou, Anshumali Shrivastava
摘要
Advancements in deep learning are often associated with increasing model sizes. Training and deploying large models require sophisticated hardware and incur significantly higher costs. Thus, model compression is a widely explored approach to solving the problem. However, SOTA techniques fall short in one or more desirable aspects of compression -for instance, pruning does not reduce memory for training, quantization can only provide up to 32× compression, Hashed-Net is cache-inefficient, etc. This paper proposes a model-agnostic, cache-friendly, and hardwareaware model compression approach: Random Operation Access Specific Tile (ROAST) hashing. ROAST collapses the parameters by clubbing them through a lightweight mapping. While clubbing these parameters, ROAST utilizes cache hierarchies by aligning the memory access pattern with the parameter access pattern. ROAST is up to ∼25× faster to train and ∼50× faster to infer than the popular parameter sharing method HashedNet. Additionally, ROAST introduces global weight sharing, which is empirically and theoretically superior to local weight sharing in HashedNet, and can be of independent interest. With ROAST, we can efficiently train and deploy the model using a much smaller memory footprint (∼ 10 -100× lesser) in text and image classification tasks. ROAST-MM kernel implementation is open-source 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- In defense of parameter sharing for model-compressionAditya Desai, Anshumali ShrivastavaICLR 2024 · 被引用 8 次
- SS1: Accelerating Inference with Fast and Expressive Sketch Structured TransformAditya Desai, Kimia Saedi, Apoorv Walia, Jihyeong Lee 等NeurIPS 2024 · 被引用 1 次
- Sketch to Adapt: Fine-Tunable Sketches for Efficient LLM AdaptationTianyi Zhang, Junda Su, Aditya Desai, Oscar Wu 等ICML 2025
它引用的顶会 Paper1
相关 Paper
- Structured Multi-Hashing for Model CompressionElad Eban, Yair Movshovitz-Attias, Hao Wu, Mark Sandler 等CVPR 2020
- ALF: Autoencoder-based Low-rank Filter-sharing for Efficient Convolutional Neural NetworksAlexander Frickenstein, Manoj Rohit Vemparala, Nael Fasfous, Laura Hauenschild 等DAC 2020 · 被引用 5 次
- DRAGONN: Distributed Randomized Approximate Gradients of Neural NetworksZhuang Wang, Zhaozhuo Xu, Xinyu Crystal Wu, Anshumali Shrivastava 等ICML 2022 · 被引用 10 次
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly 等AAAI 2021 · 被引用 79 次
- Stitchable Neural NetworksZizheng Pan, Jianfei Cai, Bohan ZhuangCVPR 2023
