MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch Training
Yang Luo, Zangwei Zheng, Ziheng Qin, Zirui Zhu, Yong Liu, Yang You
摘要
Large-batch training has become a cornerstone in accelerating the training of deep neural networks, yet it poses challenges in optimization and generalization. Existing optimizers like AdamW present performance degradation during language models' large-batch training, due to the information bottleneck in attention layers caused by the sharp increase of max attention logit. While the LAMB optimizer partially addresses this issue, some attention layers still face this issue. The reason is that l 2 -norm-based trust ratios in LAMB are less effective in directly influencing the max value of query/key weights. Furthermore, the weightwise trust ratio in LAMB is error-prone as it overlooks relationships of weight values within rows or columns. Building on these observations, we propose a novel optimizer, MERIT, which leverages the max-norm to calculate the trust ratio to constrain the max attention logit more effectively. Moreover, we further construct element-wise trust ratios to provide more robust update scaling by focusing on local weight structures. Extensive experiments of large-batch training across various sizes of GPT-2 models demonstrate the superior performance of MERIT. Notably, during the training of GPT-2 Medium, MERIT enables a 6k batch size without any performance degradation compared to the standard batch size (480) with 48B training tokens. This work highlights the importance of considering the max attention logit and finer-granularity trust ratio in large-batch training. It successfully improves the training stability and paves the way for larger batch usage, enabling faster development and iteration of large language models. Code is available at https://github.com/NUS-HPC-AI-Lab/MERIT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
相关 Paper
- Analyzing & Reducing the Need for Learning Rate Warmup in GPT TrainingAtli Kosson, Bettina Messmer, Martin JaggiNeurIPS 2024 · 被引用 25 次
- CAME: Confidence-guided Adaptive Memory Efficient OptimizationYang Luo, Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang 等ACL 2023 · 被引用 7 次
- MKOR: Momentum-Enabled Kronecker-Factor-Based Optimizer Using Rank-1 UpdatesMohammad Mozaffari, Sikan Li, Zhao Zhang, Maryam Mehri DehnaviNeurIPS 2023 · 被引用 7 次
- SLAMB: Accelerated Large Batch Training with Sparse CommunicationHang Xu, Wenxuan Zhang, Jiawei Fei, Yuzhe Wu 等ICML 2023 · 被引用 7 次
- The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-TrainingJinbo Wang, Mingze Wang, Zhanpeng Zhou, Junchi Yan 等ICML 2025
