Deeply Shared Filter Bases for Parameter-Efficient Convolutional Neural Networks
Woochul Kang, Daeyeon Kim
摘要
Modern convolutional neural networks (CNNs) have massive identical convolution blocks, and, hence, recursive sharing of parameters across these blocks has been proposed to reduce the amount of parameters. However, naive sharing of parameters poses many challenges such as limited representational power and the vanishing/exploding gradients problem of recursively shared parameters. In this paper, we present a recursive convolution block design and training method, in which a recursively shareable part, or a filter basis, is separated and learned while effectively avoiding the vanishing/exploding gradients problem during training. We show that the unwieldy vanishing/exploding gradients problem can be controlled by enforcing the elements of the filter basis orthonormal, and empirically demonstrate that the proposed orthogonality regularization improves the flow of gradients during training. Experimental results on image classification and object detection show that our approach, unlike previous parameter-sharing approaches, does not trade performance to save parameters and consistently outperforms overparameterized counterpart networks. This superior performance demonstrates that the proposed recursive convolution block design and the orthogonality regularization not only prevent performance degradation, but also consistently improve the representation capability while a significant amount of parameters are recursively shared.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Scaling-up Diverse Orthogonal Convolutional Networks by a Paraunitary FrameworkJiahao Su, Wonmin Byeon, Furong HuangICML 2022 · 被引用 1 次
- Orthogonal Convolutional Neural NetworksJiayun Wang, Yubei Chen, Rudrasis Chakraborty, Stella X. YuCVPR 2020
- Towards an Effective Orthogonal Dictionary Convolution StrategyYishi Li, Kunran Xu, Rui Lai, Lin GuAAAI 2022 · 被引用 4 次
- FSNet: Compression of Deep Convolutional Neural Networks by Filter SummaryYingzhen Yang, Jiahui Yu, Nebojsa Jojic, Jun Huan 等ICLR 2020 · 被引用 19 次
- Improved memory in recurrent neural networks with sequential non-normal dynamicsA. Emin Orhan, Xaq PitkowICLR 2020 · 被引用 16 次
