CoMERA: Computing- and Memory-Efficient Training via Rank-Adaptive Tensor Optimization
Zi Yang, Ziyue Liu, Samridhi Choudhary, Xinfeng Xie, Cao Gao, Siegfried Kunzmann, Zheng Zhang
Abstract
Training large AI models such as LLMs and DLRMs costs massive GPUs and computing time. The high training cost has become only affordable to big tech companies, meanwhile also causing increasing concerns about the environmental impact. This paper presents CoMERA, a Computing- and Memory-Efficient training method via Rank-Adaptive tensor optimization. CoMERA achieves rank-adaptive tensor-compressed (pre)-training via a multi-objective optimization formulation and improves the training to provide both a high compression ratio and excellent accuracy in the training process. Our optimized numerical computation (e.g., optimized tensorized embedding and tensor-network contractions) and GPU implementation eliminate part of the run-time overhead in the tensorized training on GPU. This leads to, for the first time, speedup per training epoch compared with standard training. CoMERA also outperforms the recent GaLore in terms of both memory and computing efficiency. Specifically, CoMERA is faster per training epoch and more memory-efficient than GaLore on a tested six-encoder transformer with single-batch training. Our method also shows speedup than standard pre-training on a BERT-like code-generation LLM while achieving compression ratio in pre-training. With further HPC optimization, CoMERA may reduce the pre-training cost of many other LLMs. An implementation of CoMERA is available at https://github.com/ziyangjoy/CoMERA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef2f4e11-aaf5-4ef8-a7ca-12afe39ca748Cited by top-tier papers4
- LaX: Boosting Low-Rank Training of Foundation Models via Latent CrossingRuijie Zhang, Ziyue Liu, Zhengyang Wang, Zheng ZhangNeurIPS 2025 · 7 citations
- DiaBlo: Diagonal Blocks Are Sufficient For FinetuningSelcuk Gurses, Aozhong Zhang, Yanxia Deng, Xun Dong et al.ICLR 2026 · 2 citations
- Calibrating and Rotating: A Unified Framework for Weight Conditioning in PEFTDa Chang, Peng Xue, Yu Li, Yongxiang Liu et al.AAAI 2026 · 2 citations
- CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank ActivationZiyue Liu, Ruijie Zhang, Zhengyang Wang, Mingsong Yan et al.EMNLP 2025
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network ModelsBeidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang et al.ICLR 2022 · 94 citations
- SVDinsTN: A Tensor Network Paradigm for Efficient Structure Search from Regularized Modeling PerspectiveYu-Bang Zheng, Xi-Le Zhao, Junhua Zeng, Chao Li et al.CVPR 2024 · 9 citations
Related papers
- Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?Xi Chen, Kaituo Feng, Changsheng Li, Xunhao Lai et al.NeurIPS 2025 · 48 citations
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 214 citations
- Tempo: Accelerating Transformer-Based Model Training through Memory Footprint ReductionMuralidhar Andoorveedu, Zhanda Zhu, Bojian Zheng, Gennady PekhimenkoNeurIPS 2022 · 8 citations
- Exploring extreme parameter compression for pre-trained language modelsBenyou Wang, Yuxin Ren, Lifeng Shang, Xin Jiang et al.ICLR 2022 · 23 citations
- On the Optimization Landscape of Low Rank Adaptation Methods for Large Language ModelsXu-Hui Liu, Yali Du, Jun Wang, Yang YuICLR 2025
