AutoShuffleNet: Learning Permutation Matrices via an Exact Lipschitz Continuous Penalty in Deep Convolutional Neural Networks
Jiancheng Lyu, Shuai Zhang, Yingyong Qi, Jack Xin
Abstract
ShuffleNet is a state-of-the-art light weight convolutional neural network architecture. Its basic operations include group, channel-wise convolution and channel shuffling. However, channel shuffling is manually designed on empirical grounds. Mathematically, shuffling is a multiplication by a permutation matrix. In this paper, we propose to automate channel shuffling by learning permutation matrices in network training. We introduce an exact Lipschitz continuous non-convex penalty so that it can be incorporated in the stochastic gradient descent to approximate permutation at high precision. Exact permutations are obtained by simple rounding at the end of training and are used in inference. The resulting network, referred to as AutoShuffleNet, achieved improved classification accuracies on data from CIFAR-10, CIFAR-100 and ImageNet while preserving the inference costs of ShuffleNet. In addition, we found experimentally that the standard convex relaxation of permutation matrices into stochastic matrices leads to poor performance. We prove theoretically the exactness (error bounds) in recovering permutation matrices when our penalty function is zero (very small). We present examples of permutation optimization through graph matching and two-layer neural network models where the loss functions are calculated in closed analytical form. In the examples, convex relaxation failed to capture permutations whereas our penalty succeeded.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60f1952f-0bca-46a8-b257-292d40d084bdCited by top-tier papers5
- Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear MapsTri Dao, Nimit Sharad Sohoni, Albert Gu, Matthew Eichhorn et al.ICLR 2020 · 55 citations
- Lipschitz Continuity Guided Knowledge DistillationYuzhang Shang, Bin Duan, Ziliang Zong, Liqiang Nie et al.ICCV 2021 · 31 citations
- Multi-Unit Transformers for Neural Machine TranslationJianhao Yan, Fandong Meng, Jie ZhouEMNLP 2020 · 21 citations
- PermLLM: Learnable Channel Permutation for N: M Sparse Large Language ModelsLancheng Zou, Shuo Yin, Zehua Pei, Tsung-Yi Ho et al.NeurIPS 2025 · 1 citation
- Spatial Assembly Networks for Image Representation LearningYang Li, Shichao Kan, Jianhe Yuan, Wenming Cao et al.CVPR 2021
Related papers
- Dynamic Region-Aware ConvolutionJin Chen, Xijun Wang, Zichao Guo, Xiangyu Zhang et al.CVPR 2021
- Lite-HRNet: A Lightweight High-Resolution NetworkChangqian Yu, Bin Xiao, Changxin Gao, Lu Yuan et al.CVPR 2021
- Convolutional Normalization: Improving Deep Convolutional Network Robustness and TrainingSheng Liu, Xiao Li, Yuexiang Zhai, Chong You et al.NeurIPS 2021 · 30 citations
- Precision Gating: Improving Neural Network Efficiency with Dynamic Dual-Precision ActivationsYichi Zhang, Ritchie Zhao, Weizhe Hua, Nayun Xu et al.ICLR 2020 · 28 citations
- Channel Permutations for N: M SparsityJeff Pool, Chong YuNeurIPS 2021 · 75 citations
