To Stay or Not to Stay in the Pre-train Basin: Insights on Ensembling in Transfer Learning
Ildus Sadrtdinov, Dmitrii Pozdeev, Dmitry P. Vetrov, Ekaterina Lobacheva
Abstract
Transfer learning and ensembling are two popular techniques for improving the performance and robustness of neural networks. Due to the high cost of pre-training, ensembles of models fine-tuned from a single pre-trained checkpoint are often used in practice. Such models end up in the same basin of the loss landscape, which we call the pre-train basin, and thus have limited diversity. In this work, we show that ensembles trained from a single pre-trained checkpoint may be improved by better exploring the pre-train basin, however, leaving the basin results in losing the benefits of transfer learning and in degradation of the ensemble quality. Based on the analysis of existing exploration methods, we propose a more effective modification of the Snapshot Ensembles (SSE) for transfer learning setup, StarSSE, which results in stronger ensembles and uniform model soups.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Merging on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model MergingAnke Tang, Enneng Yang, Li Shen, Yong Luo et al.NeurIPS 2025 · 8 citations
- Asymmetric Duos: Sidekicks Improve UncertaintyTim G. Zhou, Evan Shelhamer, Geoff PleissNeurIPS 2025 · 3 citations
- Ex Uno Pluria: Insights on Ensembling in Low Precision Number SystemsGiung Nam, Juho LeeNeurIPS 2024 · 2 citations
- The Butterfly Effect: Neural Network Training Trajectories Are Highly Sensitive to Initial ConditionsGül Sena Altintas, Devin Kwok, Colin Raffel, David RolnickICML 2025
- A Second-Order Perspective on Model Compositionality and Incremental LearningAngelo Porrello, Lorenzo Bonicelli, Pietro Buzzega, Monica Millunzi et al.ICLR 2025
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
Related papers
- TranSlider: Transfer Ensemble Learning from Exploitation to ExplorationKuo Zhong, Ying Wei, Chun Yuan, Haoli Bai et al.KDD 2020 · 12 citations
- Efficient Diversity-Driven Ensemble for Deep Neural NetworksWentao Zhang, Jiawei Jiang, Yingxia Shao, Bin CuiICDE 2020 · 17 citations
- MODEL SOUPS NEED ONLY ONE INGREDIENTAlireza Abdollahpourrostam, Nikolaos Dimitriadis, Adam Hazimeh, Pascal FrossardICML 2026
- Learning Neural Network SubspacesMitchell Wortsman, Maxwell Horton, Carlos Guestrin, Ali Farhadi et al.ICML 2021 · 101 citations
- Model ensemble instead of prompt fusion: a sample-specific knowledge transfer method for few-shot prompt tuningXiangyu Peng, Chen Xing, Prafulla Kumar Choubey, Chien-Sheng Wu et al.ICLR 2023 · 5 citations
