Exploring Mode Connectivity for Pre-trained Language Models
Yujia Qin, Cheng Qian, Jing Yi, Weize Chen, Yankai Lin, Xu Han, Zhiyuan Liu, Maosong Sun, Jie Zhou
Abstract
Recent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP. From the perspective of parameter space, PLMs provide generic initialization, starting from which high-performance minima could be found. Although plenty of works have studied how to effectively and efficiently adapt PLMs to high-performance minima, little is known about the connection of various minima reached under different adaptation configurations. In this paper, we investigate the geometric connections of different minima through the lens of mode connectivity, which measures whether two minima can be connected with a low-loss path. We conduct empirical analyses to investigate three questions: (1) how could hyperparameters, specific tuning methods, and training data affect PLM's mode connectivity? (2) How does mode connectivity change during pretraining? (3) How does the PLM's task knowledge change along the path connecting two minima? In general, exploring the mode connectivity of PLMs conduces to understanding the geometric connection of different minima, which may help us fathom the inner workings of PLM downstream adaptation. The codes are publicly available at https://github.com/ thunlp/Mode-Connectivity-PLM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e79c6006-cfe4-43bb-9a68-3f42e14a93e3Cited by top-tier papers14
- Composing Parameter-Efficient Modules with Arithmetic OperationJinghan Zhang, Shiqi Chen, Junteng Liu, Junxian HeNeurIPS 2023 · 164 citations
- Model Ratatouille: Recycling Diverse Models for Out-of-Distribution GeneralizationAlexandre Ramé, Kartik Ahuja, Jianyu Zhang, Matthieu Cord et al.ICML 2023 · 108 citations
- On the Emergence of Cross-Task Linearity in Pretraining-Finetuning ParadigmZhanpeng Zhou, Zijun Chen, Yilan Chen, Bo Zhang et al.ICML 2024 · 26 citations
- ColD Fusion: Collaborative Descent for Distributed Multitask FinetuningShachar Don-Yehiya, Elad Venezian, Colin Raffel, Noam Slonim et al.ACL 2023 · 13 citations
- Lookaround Optimizer: k steps around, 1 step averageJiangtao Zhang, Shunyu Liu, Jie Song, Tongtian Zhu et al.NeurIPS 2023 · 12 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
Related papers
- Input Space Mode Connectivity in Deep Neural NetworksJakub Vrábel, Ori Shem-Ur, Yaron Oz, David KruegerICLR 2025
- Connecting Independently Trained Modes via Layer-Wise ConnectivityYongding Tian, Zaid Al-Ars, Maksim Kitsak, H Peter HofsteeICML 2026 · 1 citation
- Fine-tuning Happens in Tiny Subspaces: Exploring Intrinsic Task-specific Subspaces of Pre-trained Language ModelsZhong Zhang, Bang Liu, Junming ShaoACL 2023 · 8 citations
- Emergence of a High-Dimensional Abstraction Phase in Language TransformersEmily Cheng, Diego Doimo, Corentin Kervadec, Iuri Macocco et al.ICLR 2025 · 1 citation
- Generalized Linear Mode Connectivity for TransformersAlexander Theus, Alessandro Cabodi, Sotiris Anagnostidis, Antonio Orvieto et al.NeurIPS 2025 · 18 citations
