AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
Di He, Songjun Tu, Ajay Jaiswal, Li Shen, Ganzhao Yuan, Shiwei Liu, Lu Yin
Abstract
Weight decay is a standard regularization technique for training large language models (LLMs). While it is common to assign a uniform decay rate to every layer, this approach overlooks the structural diversity of LLMs and the varying spectral properties across modules. In this paper, we introduce AlphaDecay, a simple yet effective method that adaptively assigns different weight decay strengths to each module of an LLM. Our approach is guided by Heavy-Tailed Self-Regularization (HT-SR) theory, which analyzes the empirical spectral density (ESD) of weight correlation matrices to quantify "heavy-tailedness." Modules exhibiting more pronounced heavy-tailed ESDs, reflecting stronger feature learning, are assigned weaker decay, while modules with lighter-tailed spectra receive stronger decay. Our method leverages tailored weight decay assignments to balance the module-wise differences in spectral properties, leading to improved performance. Extensive pre-training tasks with various model sizes from 60M to 1B demonstrate that AlphaDecay achieves better perplexity and generalization than conventional uniform decay and other adaptive decay baselines. The code is available at https://github.com/hed- ucas/AlphaDecay.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0dfcde4-80fa-44ce-820c-ca5d82171d77Cited by top-tier papers2
- RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based OptimizationShenyang Deng, Zhuoli Ouyang, Tianyu Pang, Zihang Liu et al.ICML 2026 · 7 citations
- Approaching Shannon Bound with Lossless LLM Weight CompressionHongshi Tan, Yao Chen, Gustavo Alonso, Weng-Fai Wong et al.ISCA 2026 · 2 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda et al.NeurIPS 2020 · 697 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 214 citations
- Make Your LLM Fully Utilize the ContextShengnan An, Zexiong Ma, Zeqi Lin, Nanning Zheng et al.NeurIPS 2024 · 212 citations
Related papers
- AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language ModelsHaiquan Lu, Yefan Zhou, Shiwei Liu, Zhangyang Wang et al.NeurIPS 2024 · 49 citations
- One LR Doesn’t Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMsDi He, Songjun Tu, Keyu Wang, Lu Yin et al.ICML 2026 · 3 citations
- AlphaLoRA: Assigning LoRA Experts Based on Layer Training QualityPeijun Qing, Chongyang Gao, Yefan Zhou, Xingjian Diao et al.EMNLP 2024 · 3 citations
- Weight Decay Improves Language Model PlasticityTessa Han, Sebastian Bordt, Hanlin Zhang, Sham KakadeICML 2026 · 3 citations
- Why Do We Need Weight Decay in Modern Deep Learning?Francesco D'Angelo, Maksym Andriushchenko, Aditya Vardhan Varre, Nicolas FlammarionNeurIPS 2024 · 101 citations
