HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural Networks
Jinqi Xiao, Chengming Zhang, Yu Gong, Miao Yin, Yang Sui, Lizhi Xiang, Dingwen Tao, Bo Yuan
Abstract
Low-rank compression is an important model compression strategy for obtaining compact neural network models. In general, because the rank values directly determine the model complexity and model accuracy, proper selection of layer-wise rank is very critical and desired. To date, though many low-rank compression approaches, either selecting the ranks in a manual or automatic way, have been proposed, they suffer from costly manual trials or unsatisfied compression performance. In addition, all of the existing works are not designed in a hardware-aware way, limiting the practical performance of the compressed models on real-world hardware platforms.
To address these challenges, in this paper we propose HALOC, a hardware-aware automatic low-rank compression framework. By interpreting automatic rank selection from an architecture search perspective, we develop an end-to-end solution to determine the suitable layer-wise ranks in a differentiable and hardware-aware way. We further propose design principles and mitigation strategy to efficiently explore the rank space and reduce the potential interference problem.
Experimental results on different datasets and hardware platforms demonstrate the effectiveness of our proposed approach. On CIFAR-10 dataset, HALOC enables 0.07% and 0.38% accuracy increase over the uncompressed ResNet-20 and VGG-16 models with 72.20% and 86.44% fewer FLOPs, respectively. On ImageNet dataset, HALOC achieves 0.9% higher top-1 accuracy than the original ResNet-18 model with 66.16% fewer FLOPs. HALOC also shows 0.66% higher top-1 accuracy increase than the state-of-the-art automatic low-rank compression solution with fewer computational and memory costs. In addition, HALOC demonstrates the practical speedups on different hardware platforms, verified by the measurement results on desktop GPU, embedded GPU and ASIC accelerator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8f98b71-170e-411d-a672-28a7ab19560bCited by top-tier papers3
- COMCAT: Towards Efficient Compression and Customization of Attention-Based Vision ModelsJinqi Xiao, Miao Yin, Yu Gong, Xiao Zang et al.ICML 2023 · 17 citations
- TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language ModelCheng Yang, Yang Sui, Jinqi Xiao, Lingyi Huang et al.CVPR 2025
- COAP: Memory-Efficient Training with Correlation-Aware Gradient ProjectionJinqi Xiao, Shen Sang, Tiancheng Zhi, Jing Liu et al.CVPR 2025
Builds on15
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
- SCOP: Scientific Control for Reliable Neural Network PruningYehui Tang, Yunhe Wang, Yixing Xu, Dacheng Tao et al.NeurIPS 2020 · 208 citations
- CHIP: CHannel Independence-based Pruning for Compact Neural NetworksYang Sui, Miao Yin, Yi Xie, Huy Phan et al.NeurIPS 2021 · 198 citations
- Understanding Architectures Learnt by Cell-based Neural Architecture SearchYao Shu, Wei Wang, Shaofeng CaiICLR 2020 · 92 citations
- CHEX: CHannel EXploration for CNN Model CompressionZejiang Hou, Minghai Qin, Fei Sun, Xiaolong Ma et al.CVPR 2022 · 80 citations
Related papers
- Low-Rank Compression of Neural Nets: Learning the Rank of Each LayerYerlan Idelbayev, Miguel Á. Carreira-PerpiñánCVPR 2020
- ALF: Autoencoder-based Low-rank Filter-sharing for Efficient Convolutional Neural NetworksAlexander Frickenstein, Manoj Rohit Vemparala, Nael Fasfous, Laura Hauenschild et al.DAC 2020 · 5 citations
- Compressing Neural Networks: Towards Determining the Optimal Layer-wise DecompositionLucas Liebenwein, Alaa Maalouf, Dan Feldman, Daniela RusNeurIPS 2021 · 60 citations
- Dynamical Low-Rank Compression of Neural Networks with Robustness under Adversarial AttacksSteffen Schotthöfer, Lexie Yang, Stefan SchnakeNeurIPS 2025 · 9 citations
- Neural Epitome Search for Architecture-Agnostic Network CompressionDaquan Zhou, Xiaojie Jin, Qibin Hou, Kaixin Wang et al.ICLR 2020 · 13 citations
