ExpandNets: Linear Over-parameterization to Train Compact Convolutional Networks
Shuxuan Guo, José M. Álvarez, Mathieu Salzmann
摘要
In this paper, we introduce an approach to training a given compact network. To this end, we leverage over-parameterization, which typically improves both optimization and generalization in neural network training, while being unnecessary at inference time. We propose to expand each linear layer, both fully-connected and convolutional, of the compact network into multiple linear layers, without adding any nonlinearity. As such, the resulting expanded network can benefit from over-parameterization during training but can be compressed back to the compact one algebraically at inference. We introduce several expansion strategies, together with an initialization scheme, and demonstrate the benefits of our ExpandNets on several tasks, including image classification, object detection, and semantic segmentation. As evidenced by our experiments, our approach outperforms both training the compact network from scratch and performing knowledge distillation from a teacher.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- FastViT: A Fast Hybrid Vision Transformer using Structural ReparameterizationPavan Kumar Anasosalu Vasu, James Gabriel, Jeff Zhu, Oncel Tuzel 等ICCV 2023 · 被引用 341 次
- Initialization and Regularization of Factorized Neural LayersMikhail Khodak, Neil A. Tenenholtz, Lester Mackey, Nicolò FusiICLR 2021 · 被引用 74 次
- RepSR: Training Efficient VGG-style Super-Resolution Networks with Structural Re-Parameterization and Batch NormalizationXintao Wang, Chao Dong, Ying ShanACM MM 2022 · 被引用 44 次
- Neural Collapse in Deep Linear Networks: From Balanced to Imbalanced DataHien Dang, Tho Tran Huu, Stanley J. Osher, Hung Tran-The 等ICML 2023 · 被引用 44 次
- MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced TrainingPavan Kumar Anasosalu Vasu, Hadi Pouransari, Fartash Faghri, Raviteja Vemulapalli 等CVPR 2024 · 被引用 29 次
它引用的顶会 Paper4
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution BlocksXiaohan Ding, Yuchen Guo, Guiguang Ding, Jungong HanICCV 2019 · 被引用 845 次
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park 等ICCV 2019 · 被引用 727 次
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentKarthik Abinav Sankararaman, Soham De, Zheng Xu, W. Ronny Huang 等ICML 2020 · 被引用 122 次
相关 Paper
- Beyond Student: An Asymmetric Network for Neural Network InheritanceYiyun Zhou, Jingwei Shi, Mingjing Xu, Zhonghua Jiang 等ICLR 2026 · 被引用 1 次
- Distilling Image Classifiers in Object DetectorsShuxuan Guo, José M. Álvarez, Mathieu SalzmannNeurIPS 2021 · 被引用 10 次
- Over-parameterized Student Model via Tensor Decomposition Boosted Knowledge DistillationYu-Liang Zhan, Zhong-Yi Lu, Hao Sun, Ze-Feng GaoNeurIPS 2024 · 被引用 6 次
- Distilling Knowledge via Knowledge ReviewPengguang Chen, Shu Liu, Hengshuang Zhao, Jiaya JiaCVPR 2021
- CrossKD: Cross-Head Knowledge Distillation for Object DetectionJiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li 等CVPR 2024 · 被引用 93 次
