IO-Adam: Rethinking Memory-Efficient Adaptive Optimizers from Gradient Computation
Yiting Chen, Zongwei Huo, Junchi Yan
摘要
Adaptive Moment Estimation (Adam) is one of the most popular and often the default stochastic optimizers for deep neural network training. Using first- and second-moment estimation, Adam provides adaptive learning rates for each parameter, significantly outperforming Stochastic Gradient Descent (SGD). However, as deep neural networks become larger, estimating the first and second moments consumes substantial memory. It motivates various methods to reduce memory usage for adaptive optimizers. In this paper, we propose to rethink the first and second moment estimation from a gradient computation perspective. The gradient of the weight matrix is the multiplication of the input and the gradient of the output. Instead of finding low-rank approximations of the first and second moments, as in previous work, we propose tracking the input and output gradients to efficiently estimate moments. We provide analyses of the similarities and differences between our proposed method, the widely used Adam optimizer, and previous memory-efficient optimizers designed to reduce memory usage. We conduct experiments to verify the effectiveness of our method, which reduces memory usage by up to % while preserving similar performance or even improving the performance of Adam.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 被引用 457 次
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
- LoftQ: LoRA-Fine-Tuning-aware Quantization for Large Language ModelsYixiao Li, Yifan Yu, Chen Liang, Nikos Karampatziakis 等ICLR 2024 · 被引用 217 次
- Memory Efficient Optimizers with 4-bit StatesBingrui Li, Jianfei Chen, Jun ZhuNeurIPS 2023 · 被引用 72 次
相关 Paper
- Adaptive Inertia: Disentangling the Effects of Adaptive Learning Rate and MomentumZeke Xie, Xinrui Wang, Huishuai Zhang, Issei Sato 等ICML 2022 · 被引用 65 次
- Escaping Saddle Points Faster with Stochastic MomentumJun-Kun Wang, Chi-Heng Lin, Jacob D. AbernethyICLR 2020 · 被引用 25 次
- Resetting the Optimizer in Deep RL: An Empirical StudyKavosh Asadi, Rasool Fakoor, Shoham SabachNeurIPS 2023 · 被引用 38 次
- SMMF: Square-Matricized Momentum Factorization for Memory-Efficient OptimizationKwangryeol Park, Seulki LeeAAAI 2025 · 被引用 2 次
- ACMo: Angle-Calibrated Moment Methods for Stochastic OptimizationXunpeng Huang, Runxin Xu, Hao Zhou, Zhe Wang 等AAAI 2021 · 被引用 2 次
