BAS: Bridging Adam and SignSGD for Memory-Efficient LLM Training
Yijie Zhou, Mingliang Zhang, Jiaqi Zhang, Xunliang Cai, Shi Pu
Abstract
We propose Block Adaptive Signum (BAS) , which bridges Adam and SignSGD via block-wise scaling of sign updates. By discarding element-wise second moments, BAS reduces memory overhead relative to AdamW without sacrificing performance in our tested settings. Crucially, BAS mimics Adam’s dynamics closely enough to directly inherit its hyperparameters , matching the performance of AdamW without the need for re‑tuning, a common fragility of prior low‑memory optimizers. This structural alignment makes it particularly suitable for tuning Adam-pretrained models. Furthermore, we exploit the inherent robustness of sign-based updates to store the first moment in FP8 without performance degradation. This shrinks the optimizer‑state footprint to 12.5% of AdamW’s . We theoretically prove convergence under standard assumptions and introduce a communication-efficient variant enabled by the sign-based update. Across extensive evaluations, including pre‑training a 1.5B model on 100B tokens and supervised fine-tuning of models up to 32B parameters, we demonstrate that BAS achieves performance on par with AdamW.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
- CAME: Confidence-guided Adaptive Memory Efficient OptimizationYang Luo, Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang et al.ACL 2023 · 7 citations
- From Automation to Autonomy: A Survey on Large Language Models in Scientific DiscoveryTianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang et al.EMNLP 2025 · 5 citations
- Provable Adaptivity of Adam under Non-uniform SmoothnessBohan Wang, Yushun Zhang, Huishuai Zhang, Qi Meng et al.KDD 2024 · 4 citations
Related papers
- Arbitrary-Order Block SignSGD for Memory-Efficient LLM Fine-TuningYijie Zhou, Shi PuICLR 2026
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 457 citations
- Adam-mini: Use Fewer Learning Rates To Gain MoreYushun Zhang, Congliang Chen, Ziniu Li, Tian Ding et al.ICLR 2025
- FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable TrainingPhilip Zmushko, Aleksandr Beznosikov, Martin Takác, Samuel HorváthICML 2025
- AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-trainingHuishuai Zhang, Bohan Wang, Luoxin ChenEMNLP 2025 · 1 citation
