Batch Pruning by Activation Stability
Md. Mustakin Alam, Shaker Islam, Aminul Islam
Abstract
Training deep neural networks remains costly in terms of data, time, and energy, limiting their deployment in large-scale and resource-constrained settings. To address this, we propose Batch Pruning by Activation Stability (B-PAS), a dynamic plug-in strategy that accelerates training by removing batches that contribute less to learning. B-PAS monitors the stability of activation representations across epochs and prunes batches whose activation variance exhibits minimal change, indicating diminishing learning utility. Applied to ResNet-18, ResNet-50, and the Convolutional vision Transformer (CvT) on CIFAR-10, CIFAR-100, SVHN, and ImageNet-1K, B-PAS reduces training batch usage by up to 57% with no loss in accuracy, and by 47% while slightly improving accuracy. Moreover, it achieves up to 61% savings in GPU node-hours, outperforming prior state-of-the-art pruning methods with up to 29% higher data savings and 21% greater GPU node-hour savings. We further demonstrate the generalization of B-PAS by extending it to GPT-2 fine-tuning, showing that activation stability can serve as an effective pruning signal beyond vision models. These results highlight activation stability as a powerful internal signal for efficient training, offering a practical and sustainable path toward data and energy-efficient deep learning. Our code is publicly available at https://github.com/mustakinalam/Batch-Pruning-by-Activation-Stability .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40a2008b-7fbb-4984-9a7d-98aeb8c92972Builds on19
- CvT: Introducing Convolutions to Vision TransformersHaiping Wu, Bin Xiao, Noel Codella, Mengchen Liu et al.ICCV 2021 · 2,397 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- Dataset Distillation with Infinitely Wide Convolutional NetworksTimothy Nguyen, Roman Novak, Lechao Xiao, Jaehoon LeeNeurIPS 2021 · 313 citations
Related papers
- Variance-Based Pruning for Accelerating and Compressing Trained NetworksUranik Berisha, Jens Mehnert, Alexandru Paul ConduracheICCV 2025 · 7 citations
- Winning the Lottery Ahead of Time: Efficient Early Network PruningJohn Rachwan, Daniel Zügner, Bertrand Charpentier, Simon Geisler et al.ICML 2022 · 34 citations
- Budgeted Training for Vision TransformerZhuofan Xia, Xuran Pan, Xuan Jin, Yuan He et al.ICLR 2023
- Network Expansion For Practical Training AccelerationNing Ding, Yehui Tang, Kai Han, Chao Xu et al.CVPR 2023
- InfoBatch: Lossless Training Speed Up by Unbiased Dynamic Data PruningZiheng Qin, Kai Wang, Zangwei Zheng, Jianyang Gu et al.ICLR 2024 · 94 citations
