BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language Models
Qijun Luo, Hengxu Yu, Xiao Li
摘要
This work presents BAdam, an optimization method that leverages the block coordinate descent (BCD) framework with Adam's update rule. BAdam offers a memory efficient approach to the full parameter finetuning of large language models. We conduct a theoretical convergence analysis for BAdam in the deterministic case. Experimentally, we apply BAdam to finetune the Llama 3-8B and Llama 3-70B models using a single RTX3090-24GB GPU and 4 A100-80GB GPUs, respectively. The results confirm BAdam's efficiency in terms of memory usage, running time, and optimization capability. Furthermore, the downstream performance evaluation based on MT-bench and math benchmarks shows that BAdam outperforms existing memory efficient baselines such as LoRA. It also demonstrates that BAdam can achieve comparable or even superior performance compared to Adam. Finally, the ablation study using SGD's update rule illustrates the suitability of BCD for finetuning LLMs. Our code can be easily integrated into any PyTorch-based codebase and is available at https://github.com/Ledzy/BAdam.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- EXP-Bench: Can AI Conduct AI Research Experiments?Patrick Tser Jern Kon, Qiuyi Ding, Jiachen Liu, Xinyi Zhu 等ICLR 2026 · 被引用 35 次
- AuroRA: Breaking Low-Rank Bottleneck of LoRA with Nonlinear MappingHaonan Dong, Wenhao Zhu, Guojie Song, Liang WangNeurIPS 2025 · 被引用 31 次
- SubTrack++ : Gradient Subspace Tracking for Scalable LLM TrainingSahar Rajabi, Nayeema Nonta, Sirisha RambhatlaNeurIPS 2025 · 被引用 19 次
- Harmony in Divergence: Towards Fast, Accurate, and Memory-efficient Zeroth-order LLM Fine-tuningQitao Tan, Jun Liu, Zheng Zhan, Caiwen Ding 等NeurIPS 2025 · 被引用 19 次
- Cautious Weight DecayLizhang Chen, Jonathan Li, Kaizhao Liang, Baiyu Su 等ICLR 2026 · 被引用 14 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
相关 Paper
- Accelerating Block Coordinate Descent for LLM Finetuning via Landscape ExpansionQijun Luo, Yifei Shen, Liangzu Peng, Dongsheng Li 等NeurIPS 2025 · 被引用 2 次
- Arbitrary-Order Block SignSGD for Memory-Efficient LLM Fine-TuningYijie Zhou, Shi PuICLR 2026
- Mini-batch Coresets for Memory-efficient Language Model Training on Data MixturesDang Nguyen, Wenhan Yang, Rathul Anand, Yu Yang 等ICLR 2025
- RefLoRA: Refactored Low-Rank Adaptation for Efficient Fine-Tuning of Large ModelsYilang Zhang, Bingcong Li, Georgios B. GiannakisNeurIPS 2025 · 被引用 9 次
- FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable TrainingPhilip Zmushko, Aleksandr Beznosikov, Martin Takác, Samuel HorváthICML 2025
