CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank Activation
Ziyue Liu, Ruijie Zhang, Zhengyang Wang, Mingsong Yan, Zi Yang, Paul D. Hovland, Bogdan Nicolae, Franck Cappello, Sui Tang, Zheng Zhang
Abstract
The full-size MLPs and the projection layers in attention introduce tremendous model sizes of large language models (LLMs), consuming extensive computational resources in pre-training. We empirically observe that the activations of pre-trained LLMs exhibit low-rank property. Motivated by such observations, we propose CoLA and its memory-efficient implementation, CoLA-M, to replace these full-size layers with compute-efficient auto-encoders that naturally enforce low-rank activations throughout training. This fundamental architectural change eliminates the activation redundancy and significantly boosts model capacity and training efficiency. Experiments on LLaMA models with 60 million to 7 billion parameters show that CoLA reduces the computing cost by 2× × × and improves training throughput by 1.86× × × while maintaining full-rank level performance. CoLA-M further squeezes memory cost without sacrificing throughput, offering a pre-training approach with collectively superior parameter, computing, and memory efficiency. The LLMs produced are also 2× × × smaller, enabling faster inference with lower memory cost on resource-constrained platforms. 1 * Equal contribution 1 Code available here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- AuroRA: Breaking Low-Rank Bottleneck of LoRA with Nonlinear MappingHaonan Dong, Wenhao Zhu, Guojie Song, Liang WangNeurIPS 2025 · 31 citations
- Reparameterized LLM Training via Orthogonal Equivalence TransformationZeju Qiu, Simon Buchholz, Tim Z. Xiao, Maximilian Dax et al.NeurIPS 2025 · 11 citations
- RefLoRA: Refactored Low-Rank Adaptation for Efficient Fine-Tuning of Large ModelsYilang Zhang, Bingcong Li, Georgios B. GiannakisNeurIPS 2025 · 9 citations
- CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank StructureBoao Kong, Junzhu Liang, Yuxi Liu, Renjia Deng et al.ICLR 2026 · 7 citations
- LaX: Boosting Low-Rank Training of Foundation Models via Latent CrossingRuijie Zhang, Ziyue Liu, Zhengyang Wang, Zheng ZhangNeurIPS 2025 · 7 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 214 citations
Related papers
- SLTrain: a sparse plus low rank approach for parameter and memory efficient pretrainingAndi Han, Jiaxiang Li, Wei Huang, Mingyi Hong et al.NeurIPS 2024 · 54 citations
- Scalable Efficient Training of Large Language Models with Low-dimensional Projected AttentionXingtai Lv, Ning Ding, Kaiyan Zhang, Ermo Hua et al.EMNLP 2024 · 2 citations
- Low-Rank Approximation for Sparse Attention in Multi-Modal LLMsLin Song, Yukang Chen, Shuai Yang, Xiaohan Ding et al.CVPR 2024 · 8 citations
- Parameter and Memory Efficient Pretraining via Low-rank Riemannian OptimizationZhanfeng Mo, Long-Kai Huang, Sinno Jialin PanICLR 2025
- A Token is Worth over 1, 000 Tokens: Efficient Knowledge Distillation through Low-Rank CloneJitai Hao, Qiang Huang, Hao Liu, Xinyan Xiao et al.NeurIPS 2025 · 17 citations
