Provable Meta-Learning with Low-Rank Adaptations
Jacob L. Block, Sundararajan Srinivasan, Liam Collins, Aryan Mokhtari, Sanjay Shakkottai
摘要
The power of foundation models (FMs) lies in their capacity to learn highly expressive representations that can be adapted to a broad spectrum of tasks. However, these pretrained models require additional training stages to become effective for downstream applications. In the multi-task setting, prior works have shown empirically that specific meta-learning approaches for preparing a model for future adaptation through parameter-efficient fine-tuning (PEFT) can outperform standard retraining methods, but the mechanism of the benefits of meta-learning has been largely unexplored. We introduce a framework for generic PEFT-based metalearning to learn a model that can easily adapt to unseen tasks. For linear models using LoRA, we show that standard retraining is provably suboptimal for finding an adaptable set of parameters and provide strict performance guarantees for our proposed method. We verify these theoretical insights through experiments on synthetic data as well as real-data vision and language tasks. We observe significant performance benefits using a simple implementation of our proposed meta-learning scheme during retraining relative to the conventional approach.
Lastly, although we focus on LoRA, different PEFT methods have been proposed, including variants of LoRA [25-27] and architecture adaptations [28] among others. Further, recent works have analyzed the theoretical aspects of LoRA in the fine-tuning stage [29,30], but they explored orthogonal directions to the analysis of LoRA-based meta-learning during retraining. Extended discussion of these prior works is in Appendix A.
Notation. We use bold capital letters for matrices and bold lowercase letters for vectors. N (µ, Σ) refers to the multivariate Gaussian distribution with mean µ and covariance matrix Σ. I d refers to the d × d identity matrix. ∥•∥ F refers to the Frobenius norm. S d refers to the set of d × d symmetric matrices, and S + d is the set of d × d positive semi-definite matrices. O d refers to the set of d × d orthogonal matrices.
[n] refers to the set 1, . . . , n. For a matrix X ∈ R m×n , im(X) and ker(X) refer to the image and kernel of X, while vec(X) ∈ R mn denotes the column-wise vectorization of X. For subspaces M , N , dim(M ) refers to the dimension of M and M + N = x + y|x ∈ M , y ∈ N . If M ∩ N = 0, we write the direct sum M ⊕ N .
We briefly recap the optimization process for standard retraining of an FM across multiple tasks followed by fine-tuning on a downstream task. We then introduce a general framework for PEFT-based meta-learning which adjusts the retraining phase to incorporate insights from fine-tuning.
Consider a collection of T tasks of interest T = T t T t=1 where each task T t is drawn from task distribution D and consists of n t labeled examples T t = (x t,j , y t,j ) nt j=1 . Without loss of generality, we assume consistent dimensions across tasks, so x t,j ∈ R dx , y t,j ∈ R dy for all tasks t and sample indices j. Let X t ∈ R dx×nt and Y t ∈ R dy×nt denote the concatenation of the respective input samples and labels from task t, and consider a model Φ( • ; W ) : R dx → R dy parameterized by weights W that maps feature vectors to predicted labels. We abuse notation and write Φ(X t ; W ) to denote the concatenation of Φ(x t,j ; W ) for j ∈ [n t ]. Typically W = (W 1 , . . . , W m ) is a list of matrices where W i ∈ R d×d parameterizes the i th layer of a neural network. We assume each W i is square for convenience.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
相关 Paper
- MTL-LoRA: Low-Rank Adaptation for Multi-Task LearningYaming Yang, Dilxat Muhtar, Yelong Shen, Yuefeng Zhan 等AAAI 2025 · 被引用 23 次
- LiFT: Learning to Fine-Tune via Bayesian Parameter Efficient Meta Fine-TuningMinyoung Kim, Timothy M. HospedalesICLR 2025
- Basis-Oriented Low-rank Transfer for Few-Shot and Test-Time AdaptationJunghwan Park, Woojin Cho, Junhyuk Heo, Darongsae Kwon 等CVPR 2026
- MeteoRA: Multiple-tasks Embedded LoRA for Large Language ModelsJingwei Xu, Junyu Lai, Yunpeng HuangICLR 2025
- Generalized Tensor-Based Parameter-Efficient Fine-Tuning via Lie Group TransformationsChongjie Si, Zhiyi Shi, Xuehui Wang, Yichen Xiao 等ICCV 2025
