MASS: MoErging through Adaptive Subspace Selection
Donato Crisostomi, Alessandro Zirilli, Antonio Andrea Gargiulo, Maria Sofia Bucarelli, Simone Scardapane, Fabrizio Silvestri, Iacopo Masi, Emanuele Rodolà
Abstract
Model merging has recently emerged as a lightweight alternative to ensembling, combining multiple fine-tuned models into a single set of parameters with no additional training overhead. Yet, existing merging methods fall short of matching the full accuracy of separately fine-tuned endpoints. We present MASS (MoErging through Adaptive Subspace Selection), a new approach that closes this gap by unifying multiple fine-tuned models while retaining near state-of-the-art performance across tasks. Building on the low-rank decomposition of per-task updates, MASS stores only the most salient singular components for each task and merges them into a shared model. At inference time, a non-parametric, data-free router identifies which subspace (or combination thereof) best explains an input's intermediate features and activates the corresponding task-specific block. This procedure is fully training-free and introduces only a two-pass inference overhead plus a ∼2× storage factor compared to a single pretrained model, irrespective of the number of tasks. We evaluate MASS on open-vocabulary image classification using ViT-B-16, ViT-B-32 and ViT-L-14 for benchmarks of 8, 14 and 20 tasks respectively, establishing a new state-of-the-art. Most notably, MASS recovers up to ∼ 98% of the average accuracy of individual fine-tuned models, making it a practical alternative to ensembling at a fraction of the storage cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d1334eb-eee5-43e6-a9c7-d58cc860ab57Cited by top-tier papers1
Ask how each one uses itBuilds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
Related papers
- CAT Merging: A Training-Free Approach for Resolving Conflicts in Model MergingWenju Sun, Qingyong Li, Yangliao Geng, Boyang LiICML 2025
- FW-Merging: Scaling Model Merging with Frank-Wolfe OptimizationHao Mark Chen, Shell Xu Hu, Wayne Luk, Timothy M. Hospedales et al.ICCV 2025 · 6 citations
- Unraveling LoRA Interference: Orthogonal Subspaces for Robust Model MergingHaobo Zhang, Jiayu ZhouACL 2025
- Merging on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model MergingAnke Tang, Enneng Yang, Li Shen, Yong Luo et al.NeurIPS 2025 · 8 citations
- AdaMSS: Adaptive Multi-Subspace Approach for Parameter-Efficient Fine-TuningJingjing Zheng, Wanglong Lu, Yiming Dong, Chaojie Ji et al.NeurIPS 2025
