Q3R: Quadratic Reweighted Rank Regularizer for Effective Low-Rank Training
Ipsita Ghosh, Ethan Nguyen, Christian Kümmerle
摘要
Parameter-efficient training based on low-rank optimization has become a highly successful tool for fine-tuning large deep learning models. However, these methods often fail for low-rank pre-training, where simultaneously maintaining low-rank weight structure and optimizing the task objective remains challenging. We propose the (), which leads to a novel low-rank-inducing training strategy inspired by the Iteratively Reweighted Least Squares (IRLS) framework. is based on a quadratic regularizer term that majorizes a smoothed log-determinant rank surrogate. Unlike other low-rank training techniques, can train weight matrices to prescribed low target ranks while achieving predictive performance comparable to dense models, with small computational overhead and full compatibility with existing architectures. For example, we demonstrate a -regularized ViT-Tiny experiment where truncating the model to and of its parameters results in only minor absolute accuracy drops of and , respectively, on CIFAR-10. We confirm the efficacy of on Transformers across both vision and language tasks, including low-rank fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- DoRA: Weight-Decomposed Low-Rank AdaptationShih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov 等ICML 2024 · 被引用 820 次
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 被引用 700 次
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
相关 Paper
- Low-Rank Rescaled Vision Transformer Fine-Tuning: A Residual Design ApproachWei Dong, Xing Zhang, Bihui Chen, Dawei Yan 等CVPR 2024
- PELA: Learning Parameter-Efficient Models with Low-Rank ApproximationYangyang Guo, Guangzhi Wang, Mohan S. KankanhalliCVPR 2024
- Boosting Vanilla Lightweight Vision Transformers via Re-parameterizationZhentao Tan, Xiaodan Li, Yue Wu, Qi Chu 等ICLR 2024 · 被引用 5 次
- Efficient Adaptation of Pre-trained Vision Transformer via Householder TransformationWei Dong, Yuan Sun, Yiting Yang, Xing Zhang 等NeurIPS 2024 · 被引用 10 次
- Fine-Tuning of Transformer models with FramesHarshavardhan Adepu, Li Zhang, Sanjiv Kumar, Vikas SinghICML 2026
