Compressible Dynamics in Deep Overparameterized Low-Rank Learning & Adaptation
Can Yaras, Peng Wang, Laura Balzano, Qing Qu
Abstract
While overparameterization in machine learning models offers great benefits in terms of optimization and generalization, it also leads to increased computational requirements as model sizes grow. In this work, we show that by leveraging the inherent low-dimensional structures of data and compressible dynamics within the model parameters, we can reap the benefits of overparameterization without the computational burdens. In practice, we demonstrate the effectiveness of this approach for deep low-rank matrix completion as well as fine-tuning language models. Our approach is grounded in theoretical findings for deep overparameterized low-rank matrix recovery, where we show that the learning dynamics of each weight matrix are confined to an invariant low-dimensional subspace. Consequently, we can construct and train compact, highly compressed factorizations possessing the same benefits as their overparameterized counterparts. In the context of deep matrix completion, our technique substantially improves training efficiency while retaining the advantages of overparameterization. For language model fine-tuning, we propose a method called"Deep LoRA", which improves the existing low-rank adaptation (LoRA) technique, leading to reduced overfitting and a simplified hyperparameter setup, while maintaining comparable efficiency. We validate the effectiveness of Deep LoRA on natural language tasks, particularly when fine-tuning with limited data. Our code is available at https://github.com/cjyaras/deep-lora-transformers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Exploring Low-Dimensional Subspace in Diffusion Models for Controllable Image EditingSiyi Chen, Huijie Zhang, Minzhe Guo, Yifu Lu et al.NeurIPS 2024 · 29 citations
- SubTrack++ : Gradient Subspace Tracking for Scalable LLM TrainingSahar Rajabi, Nayeema Nonta, Sirisha RambhatlaNeurIPS 2025 · 19 citations
- Towards Understanding the Dynamics of Low-Rank AdaptationShu Ding, Yang Peng, Hangan Zhou, Xinyu Lu et al.ICML 2026 · 13 citations
- Beyond Value Functions: Single-Loop Bilevel Optimization under Flatness ConditionsLiuyuan Jiang, Quan Xiao, Lisha Chen, Tianyi ChenNeurIPS 2025 · 11 citations
- AltLoRA: Towards Better Gradient Approximation in Low-Rank Adaptation with Alternating ProjectionsXin Yu, Yujia Wang, Jinghui Chen, Lingzhou XueNeurIPS 2025 · 8 citations
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li et al.NeurIPS 2021 · 303 citations
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 155 citations
- Robust Training under Label Noise by Over-parameterizationSheng Liu, Zhihui Zhu, Qing Qu, Chong YouICML 2022 · 152 citations
Related papers
- The Expressive Power of Low-Rank AdaptationYuchen Zeng, Kangwook LeeICLR 2024 · 116 citations
- GeoLoRA: Geometric integration for parameter efficient fine-tuningSteffen Schotthöfer, Emanuele Zangrando, Gianluca Ceruti, Francesco Tudisco et al.ICLR 2025
- ScaLoRA: Optimally Scaled Low-Rank Adaptation for Efficient High-Rank Fine-TuningYilang Zhang, Xiaodong Yang, Yiwei Cai, Georgios B. GiannakisICML 2026 · 1 citation
- DenseLoRA: Dense Low-Rank Adaptation of Large Language ModelsLin Mu, Xiaoyu Wang, Li Ni, Yang Li et al.ACL 2025 · 3 citations
- Balanced LoRA: Removing Parameter Invariance to Accelerate ConvergenceValérie Castin, Kimia Nadjahi, Pierre Ablin, Gabriel PeyréICML 2026
