CAME: Confidence-guided Adaptive Memory Efficient Optimization
Yang Luo, Xiaozhe Ren, Zangwei Zheng, Zhuo Jiang, Xin Jiang, Yang You
摘要
Adaptive gradient methods, such as Adam and LAMB, have demonstrated excellent performance in the training of large language models. Nevertheless, the need for adaptivity requires maintaining second-moment estimates of the per-parameter gradients, which entails a high cost of extra memory overheads. To solve this problem, several memory-efficient optimizers (e.g., Adafactor) have been proposed to obtain a drastic reduction in auxiliary memory usage, but with a performance penalty. In this paper, we first study a confidence-guided strategy to reduce the instability of existing memory efficient optimizers. Based on this strategy, we propose CAME to simultaneously achieve two goals: fast convergence as in traditional adaptive methods, and low memory usage as in memory-efficient methods. Extensive experiments demonstrate the training stability and superior performance of CAME across various NLP tasks such as BERT and GPT-2 training. Notably, for BERT pre-training on the large batch size of 32,768, our proposed optimizer attains faster convergence and higher accuracy compared with the Adam optimizer. The implementation of CAME is publicly available 1 . * Work was done when Yang Luo was an intern at Huawei Noah's Ark Lab.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 被引用 92 次
- MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable ConvergenceIonut-Vlad Modoranu, Mher Safaryan, Grigory Malinovsky, Eldar Kurtic 等NeurIPS 2024 · 被引用 32 次
- Breaking the Frozen Subspace: Importance Sampling for Low-Rank Optimization in LLM PretrainingHaochen Zhang, Junze Yin, Guanchu Wang, Zirui Liu 等NeurIPS 2025 · 被引用 7 次
- Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context ReflectionShufan Li, Konstantinos Kallidromitis, Akash Gokul, Arsh Koneru 等ICCV 2025 · 被引用 3 次
- SMMF: Square-Matricized Momentum Factorization for Memory-Efficient OptimizationKwangryeol Park, Seulki LeeAAAI 2025 · 被引用 2 次
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda 等NeurIPS 2020 · 被引用 697 次
- Towards Efficient and Scalable Sharpness-Aware MinimizationYong Liu, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh 等CVPR 2022 · 被引用 61 次
- Large Batch Optimization for Deep Learning Using New Complete Layer-Wise Adaptive Rate ScalingZhouyuan Huo, Bin Gu, Heng HuangAAAI 2021 · 被引用 35 次
相关 Paper
- FOAM: Blocked State Folding for Memory-Efficient LLM TrainingZiqing Wen, Jiahuan Wang, ping luo, Dongsheng Li 等ICML 2026 · 被引用 2 次
- SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM TrainingTianjin Huang, Ziquan Zhu, Gaojie Jin, Lu Liu 等ICLR 2025
- AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-trainingHuishuai Zhang, Bohan Wang, Luoxin ChenEMNLP 2025 · 被引用 1 次
- LDAdam: Adaptive Optimization from Low-Dimensional Gradient StatisticsThomas Robert, Mher Safaryan, Ionut-Vlad Modoranu, Dan AlistarhICLR 2025
- MERIT: Maximum-normalized Element-wise Ratio for Language Model Large-batch TrainingYang Luo, Zangwei Zheng, Ziheng Qin, Zirui Zhu 等ICML 2025
