Re-parameterizing Your Optimizers rather than Architectures
Xiaohan Ding, Honghao Chen, Xiangyu Zhang, Kaiqi Huang, Jungong Han, Guiguang Ding
Abstract
The well-designed structures in neural networks reflect the prior knowledge incorporated into the models. However, though different models have various priors, we are used to training them with model-agnostic optimizers such as SGD. In this paper, we propose to incorporate model-specific prior knowledge into optimizers by modifying the gradients according to a set of model-specific hyper-parameters. Such a methodology is referred to as Gradient Re-parameterization, and the optimizers are named RepOptimizers. For the extreme simplicity of model structure, we focus on a VGG-style plain model and showcase that such a simple model trained with a RepOptimizer, which is referred to as RepOpt-VGG, performs on par with or better than the recent well-designed models. From a practical perspective, RepOpt-VGG is a favorable base model because of its simple structure, high inference speed and training efficiency. Compared to Structural Re-parameterization, which adds priors into models via constructing extra training-time structures, RepOptimizers require no extra forward/backward computations and solve the problem of quantization. We hope to spark further research beyond the realms of model structure design. Code and models https://github.com/DingXiaoH/RepOptimizers . * Equal contributions. This work was partly done during their internships at MEGVII Technology. † Project leader. ‡ Corresponding author. 1 Prior knowledge refers to all information about the problem and the training data Krupka & Tishby (2007). Since we have not encountered any data sample while designing the model, the structural designs can be regarded as some inductive biases Mitchell (1980), which reflect our prior knowledge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da4cfdf5-3538-4abc-bdb1-48e58698ceadCited by top-tier papers13
- 3D UX-Net: A Large Kernel Volumetric ConvNet Modernizing Hierarchical Transformer for Medical Image SegmentationHo Hin Lee, Shunxing Bao, Yuankai Huo, Bennett A. LandmanICLR 2023 · 100 citations
- Make RepVGG Greater Again: A Quantization-Aware ApproachXiangxiang Chu, Liang Li, Bo ZhangAAAI 2024 · 70 citations
- Simplifying Transformer BlocksBobby He, Thomas HofmannICLR 2024 · 52 citations
- SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch NormalizationJialong Guo, Xinghao Chen, Yehui Tang, Yunhe WangICML 2024 · 40 citations
- Improved Implicit Neural Representation with Fourier Reparameterized TrainingKexuan Shi, Xingyu Zhou, Shuhang GuCVPR 2024 · 14 citations
Builds on6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 77 citations
- RepMLPNet: Hierarchical Vision MLP with Re-parameterized LocalityXiaohan Ding, Honghao Chen, Xiangyu Zhang, Jungong Han et al.CVPR 2022 · 71 citations
Related papers
- RepVGG: Making VGG-Style ConvNets Great AgainXiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han et al.CVPR 2021
- DyRep: Bootstrapping Training with Dynamic Re-parameterizationTao Huang, Shan You, Bohan Zhang, Yuxuan Du et al.CVPR 2022 · 25 citations
- Online Convolutional ReparameterizationMu Hu, Junyi Feng, Jiashen Hua, Baisheng Lai et al.CVPR 2022 · 90 citations
- RepSR: Training Efficient VGG-style Super-Resolution Networks with Structural Re-Parameterization and Batch NormalizationXintao Wang, Chao Dong, Ying ShanACM MM 2022 · 44 citations
- RIFormer: Keep Your Vision Backbone Effective But Removing Token MixerJiahao Wang, Songyang Zhang, Yong Liu, Taiqiang Wu et al.CVPR 2023
