Transporting Task Vectors across Different Architectures without Training
Filippo Rinaldi, Aniello Panariello, Giacomo Salici, Angelo Porrello, Simone Calderara
Abstract
Adapting large pre-trained models to downstream tasks often produces task-specific parameter updates that are expensive to relearn for every model variant. While recent work has shown that such updates can be transferred between models with identical architectures, transferring them across models of different widths remains unexplored. In this work, we introduce THESEUS, a trainingfree method for transporting task updates across heterogeneous-width models. Rather than matching parameters, we characterize a task update by the functional effect it induces on intermediate representations. We formalize task-vector transport as a functional matching problem on observed activations and show that, after aligning representation spaces via orthogonal Procrustes analysis, it admits a stable closed-form solution that preserves the geometry of the update. We evaluate THESEUS on vision and language models across different widths, showing consistent improvements over baselines without additional training or backpropagation. Our results show that task updates can be meaningfully transferred across architectures when task identity is defined functionally rather than parametrically. Code is available at https://github.com/ apanariello4/merge-and-rebase .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e595b3e-5ede-4fce-af05-509534c2d61fCited by top-tier papers1
Ask how each one uses itBuilds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta et al.NeurIPS 2022 · 1,483 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
Related papers
- Gradient-Sign Masking for Task Vector Transport Across Pre-Trained ModelsFilippo Rinaldi, Aniello Panariello, Giacomo Salici, Fengyuan Liu et al.ICLR 2026 · 3 citations
- Update Your Transformer to the Latest Release: Re-Basin of Task VectorsFilippo Rinaldi, Giacomo Capitani, Lorenzo Bonicelli, Donato Crisostomi et al.ICML 2025
- Fine-Tune Once, Reuse Across Models: Bayesian Task-Update Factors and ApproximationsSiyang Guo, Junbo Wang, Zibin ZhengICML 2026
- Modeling Multi-Task Model Merging as Adaptive Projective Gradient DescentYongxian Wei, Anke Tang, Li Shen, Zixuan Hu et al.ICML 2025
- Scalable Transfer Learning with Expert ModelsJoan Puigcerver, Carlos Riquelme Ruiz, Basil Mustafa, Cédric Renggli et al.ICLR 2021 · 70 citations
