BATUDE: Budget-Aware Neural Network Compression Based on Tucker Decomposition
Miao Yin, Huy Phan, Xiao Zang, Siyu Liao, Bo Yuan
Abstract
Model compression is very important for the efficient deployment of deep neural network (DNN) models on resource-constrained devices. Among various model compression approaches, high-order tensor decomposition is particularly attractive and useful because the decomposed model is very small and fully structured. For this category of approaches, tensor ranks are the most important hyper-parameters that directly determine the architecture and task performance of the compressed DNN models. However, as an NP-hard problem, selecting optimal tensor ranks under the desired budget is very challenging and the state-of-the-art studies suffer from unsatisfied compression performance and timing-consuming search procedures. To systematically address this fundamental problem, in this paper we propose BATUDE, a Budget-Aware TUcker DEcomposition-based compression approach that can efficiently calculate optimal tensor ranks via one-shot training. By integrating the rank selecting procedure to the DNN training process with a specified compression budget, the tensor ranks of the DNN models are learned from the data and thereby bringing very significant improvement on both compression ratio and classification accuracy for the compressed models. The experimental results on ImageNet dataset show that our method enjoys 0.33% top-5 higher accuracy with 2.52X less computational cost as compared to the uncompressed ResNet-18 model. For ResNet-50, the proposed approach enables 0.37% and 0.55% top-5 accuracy increase with 2.97X and 2.04X computational cost reduction, respectively, over the uncompressed model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7124587d-9511-4c71-a8eb-14d0cb7284e5Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li et al.ICCV 2019 · 540 citations
- CHIP: CHannel Independence-based Pruning for Compact Neural NetworksYang Sui, Miao Yin, Yi Xie, Huy Phan et al.NeurIPS 2021 · 198 citations
- HRank: Filter Pruning Using High-Rank Feature MapMingbao Lin, Rongrong Ji, Yan Wang, Yichen Zhang et al.CVPR 2020
- Towards Efficient Tensor Decomposition-Based DNN Model Compression With Optimization FrameworkMiao Yin, Yang Sui, Siyu Liao, Bo YuanCVPR 2021
Related papers
- Towards Compact CNNs via Collaborative CompressionYuchao Li, Shaohui Lin, Jianzhuang Liu, Qixiang Ye et al.CVPR 2021
- TDC: Towards Extremely Efficient CNNs on GPUs via Hardware-Aware Tucker DecompositionLizhi Xiang, Miao Yin, Chengming Zhang, Aravind Sukumaran-Rajam et al.PPoPP 2023 · 4 citations
- HODEC: Towards Efficient High-Order DEcomposed Convolutional Neural NetworksMiao Yin, Yang Sui, Wanzhao Yang, Xiao Zang et al.CVPR 2022 · 17 citations
- Geometry-aware training of factorized layers in tensor Tucker formatEmanuele Zangrando, Steffen Schotthöfer, Gianluca Ceruti, Jonas Kusch et al.NeurIPS 2024 · 20 citations
- Towards Extremely Compact RNNs for Video Recognition With Fully Decomposed Hierarchical Tucker StructureMiao Yin, Siyu Liao, Xiao-Yang Liu, Xiaodong Wang et al.CVPR 2021
