Deeply Shared Filter Bases for Parameter-Efficient Convolutional Neural Networks
Woochul Kang, Daeyeon Kim
Abstract
Modern convolutional neural networks (CNNs) have massive identical convolution blocks, and, hence, recursive sharing of parameters across these blocks has been proposed to reduce the amount of parameters. However, naive sharing of parameters poses many challenges such as limited representational power and the vanishing/exploding gradients problem of recursively shared parameters. In this paper, we present a recursive convolution block design and training method, in which a recursively shareable part, or a filter basis, is separated and learned while effectively avoiding the vanishing/exploding gradients problem during training. We show that the unwieldy vanishing/exploding gradients problem can be controlled by enforcing the elements of the filter basis orthonormal, and empirically demonstrate that the proposed orthogonality regularization improves the flow of gradients during training. Experimental results on image classification and object detection show that our approach, unlike previous parameter-sharing approaches, does not trade performance to save parameters and consistently outperforms overparameterized counterpart networks. This superior performance demonstrates that the proposed recursive convolution block design and the orthogonality regularization not only prevent performance degradation, but also consistently improve the representation capability while a significant amount of parameters are recursively shared.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2cb57d5-68a2-4d8f-97ff-b921398dfac9Builds on2
Related papers
- Scaling-up Diverse Orthogonal Convolutional Networks by a Paraunitary FrameworkJiahao Su, Wonmin Byeon, Furong HuangICML 2022 · 1 citation
- Orthogonal Convolutional Neural NetworksJiayun Wang, Yubei Chen, Rudrasis Chakraborty, Stella X. YuCVPR 2020
- Towards an Effective Orthogonal Dictionary Convolution StrategyYishi Li, Kunran Xu, Rui Lai, Lin GuAAAI 2022 · 4 citations
- FSNet: Compression of Deep Convolutional Neural Networks by Filter SummaryYingzhen Yang, Jiahui Yu, Nebojsa Jojic, Jun Huan et al.ICLR 2020 · 19 citations
- Improved memory in recurrent neural networks with sequential non-normal dynamicsA. Emin Orhan, Xaq PitkowICLR 2020 · 16 citations
