CR-Net: Scaling Parameter-Efficient Training with Cross-Layer Low-Rank Structure
Boao Kong, Junzhu Liang, Yuxi Liu, Renjia Deng, Kun Yuan
Abstract
Low-rank architectures have become increasingly important for efficient large language model (LLM) pre-training, providing substantial reductions in both parameter complexity and memory/computational demands. Despite these advantages, current low-rank methods face three critical shortcomings: (1) compromised model performance, (2) considerable computational overhead, and (3) limited activation memory savings. To address these limitations, we propose Cross-layer Low-Rank residual Network (CR-Net), an innovative parameter-efficient framework inspired by our discovery that inter-layer activation residuals possess low-rank properties. CR-Net implements this insight through a dual-path architecture that efficiently reconstructs layer activations by combining previous-layer outputs with their low-rank differences, thereby maintaining high-rank information with minimal parameters. We further develop a specialized activation recomputation strategy tailored for CR-Net that dramatically reduces memory requirements. Extensive pre-training experiments across model scales from 60M to 7B parameters demonstrate that CR-Net consistently outperforms state-of-the-art low-rank frameworks while requiring fewer computational resources and less memory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Synergistic Intra- and Cross-Layer Regularization Losses for MoE Expert SpecializationRizhen Hu, Yuan Cao, Boao Kong, Mou Sun et al.ICML 2026
- Attention Projection Mixing with Exogenous AnchorsJonathan SuICML 2026
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
Related papers
- CoLA: Compute-Efficient Pre-Training of LLMs via Low-Rank ActivationZiyue Liu, Ruijie Zhang, Zhengyang Wang, Mingsong Yan et al.EMNLP 2025
- PELA: Learning Parameter-Efficient Models with Low-Rank ApproximationYangyang Guo, Guangzhi Wang, Mohan S. KankanhalliCVPR 2024
- SLTrain: a sparse plus low rank approach for parameter and memory efficient pretrainingAndi Han, Jiaxiang Li, Wei Huang, Mingyi Hong et al.NeurIPS 2024 · 54 citations
- Scalable Efficient Training of Large Language Models with Low-dimensional Projected AttentionXingtai Lv, Ning Ding, Kaiyan Zhang, Ermo Hua et al.EMNLP 2024 · 2 citations
- Parameter and Memory Efficient Pretraining via Low-rank Riemannian OptimizationZhanfeng Mo, Long-Kai Huang, Sinno Jialin PanICLR 2025
