From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications
Ajay Kumar Jaiswal, Yifan Wang, Lu Yin, Shiwei Liu, Runjin Chen, Jiawei Zhao, Ananth Grama, Yuandong Tian, Zhangyang Wang
Abstract
Large Language Models' (LLMs) weight matrices can often be expressed in low-rank form with potential to relax memory and compute resource requirements. Unlike prior efforts that focus on developing novel matrix decompositions, in this work we study the non-uniform low-rank properties of weight matrices in LLMs through the lens of stabilizing gradient subspace. First, we provide a theoretical framework to understand the stabilization of gradient subspaces through Hessian analysis. Second, we empirically establish an important relationship between gradient dynamics and low-rank expressiveness of weight matrices. Our findings reveal that different LLM components exhibit varying levels of converged lowrank structures, necessitating variable rank reduction across them to minimize drop in performance due to compression. Drawing on this result, we present Weight Low-Rank Projection (WeLore) that unifies weight compression and memoryefficient fine-tuning into one, in a data-agnostic and one-shot manner. When used as a compression technique, WeLore categorizes weight matrices into Low-rank Components (LRCs) and Non-Low-rank Components (N-LRCs) and suitably encodes them for minimum performance loss. Our gradient dynamics perspective illustrates that LRCs tend to have better finetuning capabilities and their standalone finetuning can closely mimic and sometimes outperform the training loss trajectory and performance of full-finetuning with notable memory and compute footprint reduction. Codes are available at https://github. com/VITA-Group/WeLore .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ebe5ac6-095e-4004-b6e8-e8c446ece147Cited by top-tier papers7
- Flow Caching for Autoregressive Video GenerationYuexiao Ma, Xuzhe Zheng, Jing Xu, Xiwei Xu et al.ICLR 2026 · 20 citations
- LittleBit: Ultra Low-Bit Quantization via Latent FactorizationBanseok Lee, Dongkyu Kim, Youngcheon You, Youngmin KimNeurIPS 2025 · 14 citations
- Deep Hierarchical Learning with Nested Subspace Networks for Large Language ModelsPaulius Rauba, Mihaela van der SchaarICLR 2026 · 3 citations
- MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for TransformersAjay Jaiswal, Lauren Hannah, Han-Byul Kim, Duc Hoang et al.ICML 2026 · 3 citations
- POET-X: Memory-efficient LLM Training by Scaling Orthogonal TransformationZeju Qiu, Lixin LIU, Adrian Weller, Han Shi et al.ICML 2026 · 2 citations
Builds on12
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- ReLoRA: High-Rank Training Through Low-Rank UpdatesVladislav Lialin, Sherin Muckatira, Namrata Shivagunde, Anna RumshiskyICLR 2024 · 214 citations
- Language model compression with weighted low-rank factorizationYen-Chang Hsu, Ting Hua, Sungen Chang, Qian Lou et al.ICLR 2022 · 210 citations
Related papers
- SEPARATE: A Simple Low-rank Projection for Gradient Compression in Modern Large-scale Model Training ProcessHanzhen Zhao, Xingyu Xie, Cong Fang, Zhouchen LinICLR 2025
- Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?Xi Chen, Kaituo Feng, Changsheng Li, Xunhao Lai et al.NeurIPS 2025 · 48 citations
- On the Optimization Landscape of Low Rank Adaptation Methods for Large Language ModelsXu-Hui Liu, Yali Du, Jun Wang, Yang YuICLR 2025
- AdaRankGrad: Adaptive Gradient Rank and Moments for Memory-Efficient LLMs Training and Fine-TuningYehonathan Refael, Jonathan Svirsky, Boris Shustin, Wasim Huleihel et al.ICLR 2025
- Subspace Optimization for Large Language Models with Convergence GuaranteesYutong He, Pengrui Li, Yipeng Hu, Chuyan Chen et al.ICML 2025
