Fine-Tune Once, Reuse Across Models: Bayesian Task-Update Factors and Approximations
Siyang Guo, Junbo Wang, Zibin Zheng
摘要
As pre-trained models evolve rapidly, transferring fine-tuning knowledge to updated models without retraining has become a critical challenge. Most existing methods reuse parameter updates, yet the same dataset can induce substantially different updates across base models due to mismatched local loss landscapes, making such transfer unstable. We instead adopt a Bayesian-updating perspective: a base model defines a prior, while fine-tuning contributes a task-update factor that is prior-agnostic, thereby making it feasible to reuse the update across base models. Specifically, we formalize a reusable task-update factor by requiring invariance across base models and a fixed-dimensional parameterization . Our main theoretical result shows that such reusable factors exist when the variational family is a half-space, and it is already maximal among convex families. In particular, an ideal regime arises when the priors and their Bayesian posteriors remain within a shared exponential family, as it always admits a reusable update factor. Building on this existence, we propose ***B ayesian Task Update Transfer ( BTransfer ), which extracts a reusable task-update factor from a single fine-tuning run and applies it to a new prior. For deep networks, we implement BTransfer with a ``lift–transfer–return'' pipeline: 1) lift model parameters to distributions; 2) transfer the extracted task-update factor in the exponential family distributions; and 3) return the updated posterior distribution to parameter space. Extensive experiments demonstrate that our approach effectively reuses fine-tuning knowledge across models without post-training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper51
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
相关 Paper
- Update Your Transformer to the Latest Release: Re-Basin of Task VectorsFilippo Rinaldi, Giacomo Capitani, Lorenzo Bonicelli, Donato Crisostomi 等ICML 2025
- Pre-Train Your Loss: Easy Bayesian Transfer Learning with Informative PriorsRavid Shwartz-Ziv, Micah Goldblum, Hossein Souri, Sanyam Kapoor 等NeurIPS 2022 · 被引用 52 次
- Deep Reference Priors: What is the best way to pretrain a model?Yansong Gao, Rahul Ramesh, Pratik ChaudhariICML 2022 · 被引用 6 次
- A General Class of Transfer Learning Regression without Implementation CostShunya Minami, Song Liu, Stephen Wu, Kenji Fukumizu 等AAAI 2021 · 被引用 8 次
- Transporting Task Vectors across Different Architectures without TrainingFilippo Rinaldi, Aniello Panariello, Giacomo Salici, Angelo Porrello 等ICML 2026 · 被引用 3 次
