Parameter and Memory Efficient Pretraining via Low-rank Riemannian Optimization
Zhanfeng Mo, Long-Kai Huang, Sinno Jialin Pan
摘要
Pretraining large language models often requires significant computational resources and memory due to their vast parameter amount. An effective approach to enhance parameter efficiency in both training and inference is to parameterize each full-size weight as the product of two trainable low-rank factors. While lowrank fine-tuning has achieved great success, low-rank pretraining remains challenging as it requires learning extensive knowledge from scratch under the restrictive low-rank parameterization. During standard low-rank pretraining, separately optimizing the low-rank factors introduces redundant information from the full gradient, which hinders the learning process. To achieve efficient yet effective low-rank pretraining, we propose a Low-rank Riemannian Optimizer (LORO). At each LORO update step, the low-rank factor pairs are jointly updated to ensure their full-size product moves along the steepest descent direction on the low-rank manifold, without the need to compute any memory-intensive full-size matrices or gradients. Hence, our LORO finds low-rank models that achieve high performance comparable to full-size pretrained models, while significantly reducing memory usage and accelerating both training and inference. A LLaMA 1B model pretrained with LORO achieves a perplexity score of 2% better than the full-size baseline, with a 54% reduction in model memory, a ×1.8 speedup in training, and a ×2.2 speedup in inference. The code is available on GitHub 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- SubTrack++ : Gradient Subspace Tracking for Scalable LLM TrainingSahar Rajabi, Nayeema Nonta, Sirisha RambhatlaNeurIPS 2025 · 被引用 19 次
- LoRA meets Riemannion: Muon Optimizer for Parametrization-independent Low-Rank AdaptersVladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Denis Bobkov 等ICLR 2026 · 被引用 13 次
- Reparameterized LLM Training via Orthogonal Equivalence TransformationZeju Qiu, Simon Buchholz, Tim Z. Xiao, Maximilian Dax 等NeurIPS 2025 · 被引用 11 次
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank StructureBoao Kong, Junzhu Liang, Yuxi Liu, Renjia Deng 等ICLR 2026 · 被引用 7 次
- LaX: Boosting Low-Rank Training of Foundation Models via Latent CrossingRuijie Zhang, Ziyue Liu, Zhengyang Wang, Zheng ZhangNeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 被引用 457 次
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang 等ICML 2024 · 被引用 433 次
相关 Paper
- Taming Momentum: Rethinking Optimizer States Through Low-Rank ApproximationZhengbo Wang, Jian Liang, Ran He, Zilei Wang 等ICLR 2026 · 被引用 3 次
- SLTrain: a sparse plus low rank approach for parameter and memory efficient pretrainingAndi Han, Jiaxiang Li, Wei Huang, Mingyi Hong 等NeurIPS 2024 · 被引用 54 次
- AdaRankGrad: Adaptive Gradient Rank and Moments for Memory-Efficient LLMs Training and Fine-TuningYehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel 等ICLR 2025
- Train Small, Infer Large: Memory-Efficient LoRA Training for Large Language ModelsJun Zhang, Jue Wang, Huan Li, Lidan Shou 等ICLR 2025
- Riemannian Preconditioned LoRA for Fine-Tuning Foundation ModelsFangzhao Zhang, Mert PilanciICML 2024 · 被引用 43 次
