UPSCALE: Unconstrained Channel Pruning
Alvin Wan, Hanxiang Hao, Kaushik Patnaik, Yueyang Xu, Omer Hadad, David Güera, Zhile Ren, Qi Shan
Abstract
As neural networks grow in size and complexity, inference speeds decline. To combat this, one of the most effective compression techniques -- channel pruning -- removes channels from weights. However, for multi-branch segments of a model, channel removal can introduce inference-time memory copies. In turn, these copies increase inference latency -- so much so that the pruned model can be slower than the unpruned model. As a workaround, pruners conventionally constrain certain channels to be pruned together. This fully eliminates memory copies but, as we show, significantly impairs accuracy. We now have a dilemma: Remove constraints but increase latency, or add constraints and impair accuracy. In response, our insight is to reorder channels at export time, (1) reducing latency by reducing memory copies and (2) improving accuracy by removing constraints. Using this insight, we design a generic algorithm UPSCALE to prune models with any pruning pattern. By removing constraints from existing pruners, we improve ImageNet accuracy for post-training pruned models by 2.1 points on average -- benefiting DenseNet (+16.9), EfficientNetV2 (+7.9), and ResNet (+6.2). Furthermore, by reordering channels, UPSCALE improves inference speeds by up to 2x over a baseline export.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- DISP-LLM: Dimension-Independent Structural Pruning for Large Language ModelsShangqian Gao, Chi-Heng Lin, Ting Hua, Zheng Tang et al.NeurIPS 2024 · 42 citations
- RTInfer: Real-Time Inference of Multiple DNNs on Edge GPUsRenjie Li, Tong Sun, Yi Gao, Wei DongICML 2026
Builds on2
Related papers
- Group Fisher Pruning for Practical Network CompressionLiyang Liu, Shilong Zhang, Zhanghui Kuang, Aojun Zhou et al.ICML 2021 · 204 citations
- Efficient Latency-Aware CNN Depth Compression via Two-Stage Dynamic ProgrammingJinuk Kim, Yeonwoo Jeong, Deokjae Lee, Hyun Oh SongICML 2023 · 1 citation
- DFPC: Data flow driven pruning of coupled channels without dataTanay Narshana, Chaitanya Murti, Chiranjib BhattacharyyaICLR 2023
- Towards Efficient Model Compression via Learned Global RankingTing-Wu Chin, Ruizhou Ding, Cha Zhang, Diana MarculescuCVPR 2020
- Unified Data-Free Compression: Pruning and Quantization without Fine-TuningShipeng Bai, Jun Chen, Xintian Shen, Yixuan Qian et al.ICCV 2023 · 31 citations
