TVSPrune - Pruning Non-discriminative filters via Total Variation separability of intermediate representations without fine tuning
Chaitanya Murti, Tanay Narshana, Chiranjib Bhattacharyya
Abstract
Achieving structured, data-free sparsity of deep neural networks (DNNs) remains an open area of research. In this work, we address the challenge of pruning filters without access to the original training set or loss function. We propose the discriminative filters hypothesis, that well-trained models possess discriminative filters, and any non-discriminative filters can be pruned without impacting the predictive performance of the classifier. Based on this hypothesis, we propose a new paradigm for pruning neural networks: distributional pruning, wherein we only require access to the distributions that generated the original datasets. Our approach to solving the problem of formalising and quantifying the discriminating ability of filters is through the total variation (TV) distance between the class-conditional distributions of the filter outputs. We present empirical results that, using this definition of discriminability, support our hypothesis on a variety of datasets and architectures. Next, we define the LDIFF score, a heuristic to quantify the extent to which a layer possesses a mixture of discriminative and non-discriminative filters. We empirically demonstrate that the LDIFF score is indicative of the performance of random pruning for a given layer, and thereby indicates the extent to which a layer may be pruned. Our main contribution is a novel one-shot pruning algorithm, called TVSPrune, that identifies non-discriminative filters for pruning. We extend this algorithm to IterTVSPrune, wherein we iteratively apply TVSPrune, thereby enabling us to achieve greater sparsity. Last, we demonstrate the efficacy of the TVSPrune on a variety of datasets, and show that in some cases, we can prune up to 60% of parameters with only a 2% loss of accuracy without any fine-tuning of the model, beating the nearest baseline by almost 10%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- DisCEdit: Model Editing by Identifying Discriminative ComponentsChaitanya Murti, Chiranjib BhattacharyyaNeurIPS 2024 · 2 citations
- ModHiFi: Identifying High Fidelity predictive components for Model ModificationDhruva Kashyap, Chaitanya Murti, Pranav K. Nayak, Tanay Narshana et al.NeurIPS 2025 · 1 citation
Related papers
- RED : Looking for Redundancies for Data-FreeStructured Compression of Deep Neural NetworksEdouard Yvinec, Arnaud Dapogny, Matthieu Cord, Kevin BaillyNeurIPS 2021 · 32 citations
- CDP: Towards Optimal Filter Pruning via Class-wise Discriminative PowerTianshuo Xu, Yuhang Wu, Xiawu Zheng, Teng Xi et al.ACM MM 2021 · 5 citations
- Bayesian based Re-parameterization for DNN Model PruningXiaotong Lu, Teng Xi, Baopu Li, Gang Zhang et al.ACM MM 2022 · 4 citations
- Provable Filter Pruning for Efficient Neural NetworksLucas Liebenwein, Cenk Baykal, Harry Lang, Dan Feldman et al.ICLR 2020 · 161 citations
- DPFPS: Dynamic and Progressive Filter Pruning for Compressing Convolutional Neural Networks from ScratchXiaofeng Ruan, Yufan Liu, Bing Li, Chunfeng Yuan et al.AAAI 2021 · 49 citations
