Towards Efficient Tensor Decomposition-Based DNN Model Compression With Optimization Framework
Miao Yin, Yang Sui, Siyu Liao, Bo Yuan
Abstract
Advanced tensor decomposition, such as tensor train (TT) and tensor ring (TR), has been widely studied for deep neural network (DNN) model compression, especially for recurrent neural networks (RNNs). However, compressing convolutional neural networks (CNNs) using TT/TR always suffers significant accuracy loss. In this paper, we propose a systematic framework for tensor decomposition-based model compression using Alternating Direction Method of Multipliers (ADMM). By formulating TT decompositionbased model compression to an optimization problem with constraints on tensor ranks, we leverage ADMM technique to systemically solve this optimization problem in an iterative way. During this procedure, the entire DNN model is trained in the original structure instead of TT format, but gradually enjoys the desired low tensor rank characteristics. We then decompose this uncompressed model to TT format, and fine-tune it to finally obtain a high-accuracy TTformat DNN model. Our framework is very general, and it works for both CNNs and RNNs, and can be easily modified to fit other tensor decomposition approaches. We evaluate our proposed framework on different DNN models for image classification and video recognition tasks. Experimental results show that our ADMM-based TT-format models demonstrate very high compression performance with high accuracy. Notably, on CIFAR-100, with 2.3ˆand 2.4ˆcompression ratios, our models have 1.96% and 2.21% higher top-1 accuracy than the original on ImageNet, our model achieves 2.47ˆFLOPs reduction without accuracy loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97bd66f9-bbc8-4c8b-b24f-4a8819b5a4b6Cited by top-tier papers15
- CHIP: CHannel Independence-based Pruning for Compact Neural NetworksYang Sui, Miao Yin, Yi Xie, Huy Phan et al.NeurIPS 2021 · 198 citations
- Compressible-composable NeRF via Rank-residual DecompositionJiaxiang Tang, Xiaokang Chen, Jingbo Wang, Gang ZengNeurIPS 2022 · 125 citations
- Approximate Caching for Efficiently Serving Text-to-Image Diffusion ModelsShubham Agarwal, Subrata Mitra, Sarthak Chakraborty, Srikrishna Karanam et al.NSDI 2024 · 44 citations
- BATUDE: Budget-Aware Neural Network Compression Based on Tucker DecompositionMiao Yin, Huy Phan, Xiao Zang, Siyu Liao et al.AAAI 2022 · 37 citations
- HALOC: Hardware-Aware Automatic Low-Rank Compression for Compact Neural NetworksJinqi Xiao, Chengming Zhang, Yu Gong, Miao Yin et al.AAAI 2023 · 35 citations
Related papers
- Towards Extremely Compact RNNs for Video Recognition With Fully Decomposed Hierarchical Tucker StructureMiao Yin, Siyu Liao, Xiao-Yang Liu, Xiaodong Wang et al.CVPR 2021
- TDC: Towards Extremely Efficient CNNs on GPUs via Hardware-Aware Tucker DecompositionLizhi Xiang, Miao Yin, Chengming Zhang, Aravind Sukumaran-Rajam et al.PPoPP 2023 · 4 citations
- Fully-Connected Tensor Network Decomposition and Its Application to Higher-Order Tensor CompletionYu-Bang Zheng, Ting-Zhu Huang, Xi-Le Zhao, Qibin Zhao et al.AAAI 2021 · 183 citations
- Towards Compact CNNs via Collaborative CompressionYuchao Li, Shaohui Lin, Jianzhuang Liu, Qixiang Ye et al.CVPR 2021
- Tensor FISTA-Net for Real-Time Snapshot Compressive ImagingXiaochen Han, Bo Wu, Zheng Shou, Xiao-Yang Liu et al.AAAI 2020 · 46 citations
