Re-parameterizing Your Optimizers rather than Architectures
Xiaohan Ding, Honghao Chen, Xiangyu Zhang, Kaiqi Huang, Jungong Han, Guiguang Ding
摘要
The well-designed structures in neural networks reflect the prior knowledge incorporated into the models. However, though different models have various priors, we are used to training them with model-agnostic optimizers such as SGD. In this paper, we propose to incorporate model-specific prior knowledge into optimizers by modifying the gradients according to a set of model-specific hyper-parameters. Such a methodology is referred to as Gradient Re-parameterization, and the optimizers are named RepOptimizers. For the extreme simplicity of model structure, we focus on a VGG-style plain model and showcase that such a simple model trained with a RepOptimizer, which is referred to as RepOpt-VGG, performs on par with or better than the recent well-designed models. From a practical perspective, RepOpt-VGG is a favorable base model because of its simple structure, high inference speed and training efficiency. Compared to Structural Re-parameterization, which adds priors into models via constructing extra training-time structures, RepOptimizers require no extra forward/backward computations and solve the problem of quantization. We hope to spark further research beyond the realms of model structure design. Code and models https://github.com/DingXiaoH/RepOptimizers . * Equal contributions. This work was partly done during their internships at MEGVII Technology. † Project leader. ‡ Corresponding author. 1 Prior knowledge refers to all information about the problem and the training data Krupka & Tishby (2007). Since we have not encountered any data sample while designing the model, the structural designs can be regarded as some inductive biases Mitchell (1980), which reflect our prior knowledge.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- 3D UX-Net: A Large Kernel Volumetric ConvNet Modernizing Hierarchical Transformer for Medical Image SegmentationHo Hin Lee, Shunxing Bao, Yuankai Huo, Bennett A. LandmanICLR 2023 · 被引用 100 次
- Make RepVGG Greater Again: A Quantization-Aware ApproachXiangxiang Chu, Liang Li, Bo ZhangAAAI 2024 · 被引用 70 次
- Simplifying Transformer BlocksBobby He, Thomas HofmannICLR 2024 · 被引用 52 次
- SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch NormalizationJialong Guo, Xinghao Chen, Yehui Tang, Yunhe WangICML 2024 · 被引用 40 次
- Improved Implicit Neural Representation with Fourier Reparameterized TrainingKexuan Shi, Xingyu Zhou, Shuhang GuCVPR 2024 · 被引用 14 次
它引用的顶会 Paper6
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Unbiased Gradient Estimation in Unrolled Computation Graphs with Persistent Evolution StrategiesPaul Vicol, Luke Metz, Jascha Sohl-DicksteinICML 2021 · 被引用 77 次
- RepMLPNet: Hierarchical Vision MLP with Re-parameterized LocalityXiaohan Ding, Honghao Chen, Xiangyu Zhang, Jungong Han 等CVPR 2022 · 被引用 71 次
相关 Paper
- RepVGG: Making VGG-Style ConvNets Great AgainXiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han 等CVPR 2021
- DyRep: Bootstrapping Training with Dynamic Re-parameterizationTao Huang, Shan You, Bohan Zhang, Yuxuan Du 等CVPR 2022 · 被引用 25 次
- Online Convolutional ReparameterizationMu Hu, Junyi Feng, Jiashen Hua, Baisheng Lai 等CVPR 2022 · 被引用 90 次
- RepSR: Training Efficient VGG-style Super-Resolution Networks with Structural Re-Parameterization and Batch NormalizationXintao Wang, Chao Dong, Ying ShanACM MM 2022 · 被引用 44 次
- RIFormer: Keep Your Vision Backbone Effective But Removing Token MixerJiahao Wang, Songyang Zhang, Yong Liu, Taiqiang Wu 等CVPR 2023
