Lune

NeurIPS2025顶会

Provable Meta-Learning with Low-Rank Adaptations

Jacob L. Block, Sundararajan Srinivasan, Liam Collins, Aryan Mokhtari, Sanjay Shakkottai

2025年份
2被引次数

摘要

The power of foundation models (FMs) lies in their capacity to learn highly expressive representations that can be adapted to a broad spectrum of tasks. However, these pretrained models require additional training stages to become effective for downstream applications. In the multi-task setting, prior works have shown empirically that specific meta-learning approaches for preparing a model for future adaptation through parameter-efficient fine-tuning (PEFT) can outperform standard retraining methods, but the mechanism of the benefits of meta-learning has been largely unexplored. We introduce a framework for generic PEFT-based metalearning to learn a model that can easily adapt to unseen tasks. For linear models using LoRA, we show that standard retraining is provably suboptimal for finding an adaptable set of parameters and provide strict performance guarantees for our proposed method. We verify these theoretical insights through experiments on synthetic data as well as real-data vision and language tasks. We observe significant performance benefits using a simple implementation of our proposed meta-learning scheme during retraining relative to the conventional approach.

Lastly, although we focus on LoRA, different PEFT methods have been proposed, including variants of LoRA [25-27] and architecture adaptations [28] among others. Further, recent works have analyzed the theoretical aspects of LoRA in the fine-tuning stage [29,30], but they explored orthogonal directions to the analysis of LoRA-based meta-learning during retraining. Extended discussion of these prior works is in Appendix A.

Notation. We use bold capital letters for matrices and bold lowercase letters for vectors. N (µ, Σ) refers to the multivariate Gaussian distribution with mean µ and covariance matrix Σ. I d refers to the d × d identity matrix. ∥•∥ F refers to the Frobenius norm. S d refers to the set of d × d symmetric matrices, and S + d is the set of d × d positive semi-definite matrices. O d refers to the set of d × d orthogonal matrices.

[n] refers to the set 1, . . . , n. For a matrix X ∈ R m×n , im(X) and ker(X) refer to the image and kernel of X, while vec(X) ∈ R mn denotes the column-wise vectorization of X. For subspaces M , N , dim(M ) refers to the dimension of M and M + N = x + y|x ∈ M , y ∈ N . If M ∩ N = 0, we write the direct sum M ⊕ N .

We briefly recap the optimization process for standard retraining of an FM across multiple tasks followed by fine-tuning on a downstream task. We then introduce a general framework for PEFT-based meta-learning which adjusts the retraining phase to incorporate insights from fine-tuning.

Consider a collection of T tasks of interest T = T t T t=1 where each task T t is drawn from task distribution D and consists of n t labeled examples T t = (x t,j , y t,j ) nt j=1 . Without loss of generality, we assume consistent dimensions across tasks, so x t,j ∈ R dx , y t,j ∈ R dy for all tasks t and sample indices j. Let X t ∈ R dx×nt and Y t ∈ R dy×nt denote the concatenation of the respective input samples and labels from task t, and consider a model Φ( • ; W ) : R dx → R dy parameterized by weights W that maps feature vectors to predicted labels. We abuse notation and write Φ(X t ; W ) to denote the concatenation of Φ(x t,j ; W ) for j ∈ [n t ]. Typically W = (W 1 , . . . , W m ) is a list of matrices where W i ∈ R d×d parameterizes the i th layer of a neural network. We assume each W i is square for convenience.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper15

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖