SubTrack++ : Gradient Subspace Tracking for Scalable LLM Training
Sahar Rajabi, Nayeema Nonta, Sirisha Rambhatla
摘要
Training large language models (LLMs) is highly resource-intensive due to their massive number of parameters and the overhead of optimizer states. While recent work has aimed to reduce memory consumption, such efforts often entail trade-offs among memory efficiency, training time, and model performance. Yet, true democratization of LLMs requires simultaneous progress across all three dimensions. To this end, we propose SubTrack++ that leverages Grassmannian gradient subspace tracking combined with projection-aware optimizers, enabling Adam's internal statistics to adapt to subspace changes. Additionally, employing recovery scaling, a technique that restores information lost through low-rank projections, further enhances model performance. Our method demonstrates SOTA convergence by exploiting Grassmannian geometry, reducing training wall-time by up to 65% compared to the best performing baseline, LDAdam, while preserving the reduced memory footprint. Code is at https://github.com/criticalml-uw/SubTrack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 被引用 457 次
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
相关 Paper
- LDAdam: Adaptive Optimization from Low-Dimensional Gradient StatisticsThomas Robert, Mher Safaryan, Ionut-Vlad Modoranu, Dan AlistarhICLR 2025
- GRASS: Compute Efficient Low-Memory LLM Training with Structured Sparse GradientsAashiq Muhamed, Oscar Li, David P. Woodruff, Mona T. Diab 等EMNLP 2024 · 被引用 1 次
- A Memory Efficient Randomized Subspace Optimization Method for Training Large Language ModelsYiming Chen, Yuan Zhang, Yin Liu, Kun Yuan 等ICML 2025
- COSMOS: A Hybrid Adaptive Optimizer for Efficient Training of Large Language ModelsLiming Liu, Zhenghao Xu, Zixuan Zhang, Hao Kang 等ICLR 2026 · 被引用 27 次
- Lean and Mean Adaptive Optimization via Subset-Norm and Subspace-Momentum with Convergence GuaranteesThien Hang Nguyen, Huy L. NguyenICML 2025
