PLeaS - Merging Models with Permutations and Least Squares
Anshul Nasery, Jonathan Hayase, Pang Wei Koh, Sewoong Oh
Abstract
The democratization of machine learning systems has made the process of fine-tuning accessible to practitioners, leading to a wide range of open-source models fine-tuned on specialized tasks and datasets. Recent work has proposed to merge such models to combine their functionalities. However, prior approaches are usually restricted to models that are fine-tuned from the same base model. Furthermore, the final merged model is typically required to be of the same size as the original models. In this work, we propose a new two-step algorithm to merge models-termed PLeaS-which relaxes these constraints. First, leveraging the Permutation symmetries inherent in the two models, PLeaS partially matches nodes in each layer by maximizing alignment. Next, PLeaS computes the weights of the merged model as a layer-wise Least Squares solution to minimize the approximation error between the features of the merged model and the permuted features of the original models. PLeaS allows a practitioner to merge two models sharing the same architecture into a single performant model of a desired size, even when the two original models are fine-tuned from different base models. We also demonstrate how our method can be extended to address a challenging scenario where no data is available from the fine-tuning domains. We demonstrate our method to merge ResNet and ViT models trained with shared and different label spaces, and show improvement over the state-of-the-art merging methods of up to 15 percentage points for the same target compute while merging models trained on Domain-Net and fine-grained classification tasks 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cd4ee70-7d2e-4d11-aae0-638bcf214c71Cited by top-tier papers12
- Scalable Fingerprinting of Large Language ModelsAnshul Nasery, Jonathan Hayase, Creston Brooks, Peiyao Sheng et al.NeurIPS 2025 · 17 citations
- Robust Fine-tuning of Vision-Language-Action Robot Policies via Parameter MergingYajat Yadav, Zhiyuan Zhou, Andrew Wagenmaker, Karl Pertsch et al.ICLR 2026 · 10 citations
- Merging on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model MergingAnke Tang, Enneng Yang, Li Shen, Yong Luo et al.NeurIPS 2025 · 8 citations
- Transporting Task Vectors across Different Architectures without TrainingFilippo Rinaldi, Aniello Panariello, Giacomo Salici, Angelo Porrello et al.ICML 2026 · 3 citations
- Gradient-Sign Masking for Task Vector Transport Across Pre-Trained ModelsFilippo Rinaldi, Aniello Panariello, Giacomo Salici, Fengyuan Liu et al.ICLR 2026 · 3 citations
Builds on16
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang et al.ICML 2024 · 605 citations
- Model Fusion via Optimal TransportSidak Pal Singh, Martin JaggiNeurIPS 2020 · 330 citations
Related papers
- Navigating the Accuracy-Size Trade-Off with Flexible Model MergingAkash Balasaheb Dhasade, Divyansh Jhunjhunwala, Milos Vujasinovic, Gauri Joshi et al.ICLR 2026 · 2 citations
- FW-Merging: Scaling Model Merging with Frank-Wolfe OptimizationHao Mark Chen, Shell Xu Hu, Wayne Luk, Timothy M. Hospedales et al.ICCV 2025 · 6 citations
- ZipIt! Merging Models from Different Tasks without TrainingGeorge Stoica, Daniel Bolya, Jakob Bjorner, Pratik Ramesh et al.ICLR 2024 · 185 citations
- Training-free LLM Merging for Multi-task LearningZichuan Fu, Xian Wu, Yejing Wang, Wanyu Wang et al.ACL 2025
- Dataless Knowledge Fusion by Merging Weights of Language ModelsXisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, Pengxiang ChengICLR 2023 · 8 citations
