IDAP++: Advancing Divergence-Aware Pruning with Joint Filter and Layer Optimization
Aleksei Samarin, Artem A. Nazarenko, Egor Kotenko, Alexander Savelev, Aleksei Toropov, Alexandr Motyko, Valentin Malykh
摘要
Modern knowledge and large volumes of data are increasingly encoded within neural networks, making the task of simplifying their structures and reducing the number of parameters especially relevant, both to improve efficiency and to facilitate deployment in resource-constrained environments. This paper presents a novel approach to neural network compression that addresses redundancy at both the filter and architectural levels through a unified framework grounded in information flow analysis. Building upon the concept of tensor flow divergence, which quantifies how information transforms across network layers, we develop a two-stage optimization process. The first stage employs iterative divergence-aware pruning to identify and remove redundant filters while preserving critical information pathways. The second stage extends this principle to higher-level architecture optimization by analyzing layer-wise contributions to information propagation and selectively eliminating entire layers that demonstrate minimal impact on network performance. The proposed method naturally adapts to diverse architectures, including convolutional networks, transformers, and hybrid designs, providing a consistent metric for comparing the structural importance across different layer types. Experimental validation across multiple modern architectures and datasets reveals that this combined approach achieves substantial model compression while maintaining competitive accuracy. The presented approach achieves parameter reduction results that are globally comparable to state-of-the-art solutions and outperform them across a wide range of modern neural network architectures, from convolutional models to transformers. The results demonstrate how flow divergence serves as an effective guiding principle for both filter-level and layer-level optimization, offering practical benefits for deployment in resource-constrained environments.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Convolutional Neural Network Pruning With Structural Redundancy ReductionZi Wang, Chengcheng Li, Xiangyang WangCVPR 2021
- Provable Filter Pruning for Efficient Neural NetworksLucas Liebenwein, Cenk Baykal, Harry Lang, Dan Feldman 等ICLR 2020 · 被引用 161 次
- Towards Meta-Pruning via Optimal TransportAlexander Theus, Olin Geimer, Friedrich Wicke, Thomas Hofmann 等ICLR 2024 · 被引用 8 次
- Towards Compact CNNs via Collaborative CompressionYuchao Li, Shaohui Lin, Jianzhuang Liu, Qixiang Ye 等CVPR 2021
- CHIP: CHannel Independence-based Pruning for Compact Neural NetworksYang Sui, Miao Yin, Yi Xie, Huy Phan 等NeurIPS 2021 · 被引用 198 次
