GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training
Chen Zhu, Renkun Ni, Zheng Xu, Kezhi Kong, W. Ronny Huang, Tom Goldstein
Abstract
Innovations in neural architectures have fostered significant breakthroughs in language modeling and computer vision. Unfortunately, novel architectures often result in challenging hyper-parameter choices and training instability if the network parameters are not properly initialized. A number of architecture-specific initialization schemes have been proposed, but these schemes are not always portable to new architectures. This paper presents GradInit, an automated and architecture agnostic method for initializing neural networks. GradInit is based on a simple heuristic; the norm of each network layer is adjusted so that a single step of SGD or Adam with prescribed hyperparameters results in the smallest possible loss value. This adjustment is done by introducing a scalar multiplier variable in front of each parameter block, and then optimizing these variables using a simple numerical scheme. GradInit accelerates the convergence and test performance of many convolutional architectures, both with or without skip connections, and even without normalization layers. It also improves the stability of the original Transformer architecture for machine translation, enabling training it without learning rate warmup using either Adam or SGD under a wide range of learning rates and momentum coefficients. Code is available at https://github.com/zhuchen03/gradinit .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dab0134d-4ebb-40f7-b7c6-6c98761e795fCited by top-tier papers21
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Cramming: Training a Language Model on a single GPU in one dayJonas Geiping, Tom GoldsteinICML 2023 · 115 citations
- Hyper-Representations as Generative Models: Sampling Unseen Neural Network WeightsKonstantin Schürholt, Boris Knyazev, Xavier Giró-i-Nieto, Damian BorthNeurIPS 2022 · 78 citations
- No Train No Gain: Revisiting Efficient Training Algorithms For Transformer-based Language ModelsJean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini et al.NeurIPS 2023 · 63 citations
- The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of TransformersRóbert Csordás, Kazuki Irie, Jürgen SchmidhuberEMNLP 2021 · 55 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
Related papers
- Improving Transformer Optimization Through Better InitializationXiao Shi Huang, Felipe Pérez, Jimmy Ba, Maksims VolkovsICML 2020 · 181 citations
- Towards Theoretically Inspired Neural Initialization OptimizationYibo Yang, Hong Wang, Haobo Yuan, Zhouchen LinNeurIPS 2022 · 15 citations
- AutoInit: Analytic Signal-Preserving Weight Initialization for Neural NetworksGarrett Bingham, Risto MiikkulainenAAAI 2023 · 6 citations
- Principled Architecture-aware Scaling of HyperparametersWuyang Chen, Junru Wu, Zhangyang Wang, Boris HaninICLR 2024 · 3 citations
- When Will Gradient Regularization Be Harmful?Yang Zhao, Hao Zhang, Xiuyuan HuICML 2024 · 3 citations
