Locally Free Weight Sharing for Network Width Search
Xiu Su, Shan You, Tao Huang, Fei Wang, Chen Qian, Changshui Zhang, Chang Xu
Abstract
Searching for network width is an effective way to slim deep neural networks with hardware budgets. With this aim, a one-shot supernet is usually leveraged as a performance evaluator to rank the performance w.r.t. different width. Nevertheless, current methods mainly follow a manually fixed weight sharing pattern, which is limited to distinguish the performance gap of different width. In this paper, to better evaluate each width, we propose a loCAlly FrEe weight sharing strategy (CafeNet) accordingly. In CafeNet, weights are more freely shared, and each width is jointly indicated by its base channels and free channels, where free channels are supposed to locate freely in a local zone to better represent each width. Besides, we propose to further reduce the search space by leveraging our introduced FLOPs-sensitive bins. As a result, our CafeNet can be trained stochastically and get optimized within a min-min strategy. Extensive experiments on ImageNet, CIFAR-10, CelebA and MS COCO dataset have verified our superiority comparing to other state-of-the-art baselines. For example, our method can further boost the benchmark NAS network EfficientNet-B0 by 0.41% via searching its width more delicately. * Corresponding author. 1 Other literature also use the number of channels/filters to indicate the network width.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8fa38dc0-9ae5-4f3c-a62b-66f9accbc659Cited by top-tier papers23
- Patch Slimming for Efficient Vision TransformersYehui Tang, Kai Han, Yunhe Wang, Chang Xu et al.CVPR 2022 · 173 citations
- CHEX: CHannel EXploration for CNN Model CompressionZejiang Hou, Minghai Qin, Fei Sun, Xiaolong Ma et al.CVPR 2022 · 80 citations
- SOSP: Efficiently Capturing Global Correlations by Second-Order Structured PruningManuel Nonnenmacher, Thomas Pfeil, Ingo Steinwart, David ReebICLR 2022 · 48 citations
- K-shot NAS: Learnable Weight-Sharing for NAS with K-shot SupernetsXiu Su, Shan You, Mingkai Zheng, Fei Wang et al.ICML 2021 · 38 citations
- Dynamic Sparse Training with Structured SparsityMike Lasby, Anna Golubeva, Utku Evci, Mihai Nica et al.ICLR 2024 · 37 citations
Builds on9
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- MetaPruning: Meta Learning for Automatic Neural Network Channel PruningZechun Liu, Haoyuan Mu, Xiangyu Zhang, Zichao Guo et al.ICCV 2019 · 633 citations
- Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient SpaceShangchen Du, Shan You, Xiaojie Li, Jianlong Wu et al.NeurIPS 2020 · 144 citations
- Good Subnetworks Provably Exist: Pruning via Greedy Forward SelectionMao Ye, Chengyue Gong, Lizhen Nie, Denny Zhou et al.ICML 2020 · 123 citations
- ISTA-NAS: Efficient and Consistent Neural Architecture Search by Sparse CodingYibo Yang, Hongyang Li, Shan You, Fei Wang et al.NeurIPS 2020 · 66 citations
Related papers
- BCNet: Searching for Network Width With Bilaterally Coupled NetworkXiu Su, Shan You, Fei Wang, Chen Qian et al.CVPR 2021
- Distribution Consistent Neural Architecture SearchJunyi Pan, Chong Sun, Yizhou Zhou, Ying Zhang et al.CVPR 2022 · 9 citations
- AlphaNet: Improved Training of Supernets with Alpha-DivergenceDilin Wang, Chengyue Gong, Meng Li, Qiang Liu et al.ICML 2021 · 52 citations
- Few-Shot Neural Architecture SearchYiyang Zhao, Linnan Wang, Yuandong Tian, Rodrigo Fonseca et al.ICML 2021 · 100 citations
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 444 citations
