Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training
WenJie Zhou, Bohan Wang, Hongtao Zhang, Chenxi Jia, Wei Chen, Xueqi Cheng
Abstract
Model merging has emerged as a lightweight paradigm for enhancing Large Language Models (LLMs), yet its underlying mechanisms remain poorly understood. In this work, we analyze late-stage pre-training trajectories and uncover a Rank-1 Subspace phenomenon: while raw optimization steps oscillate violently, consecutive merged checkpoints collapse onto a stable, approximately one-dimensional linear manifold. We theoretically ground this observation in a river-valley landscape analysis: averaging acts as a geometric low-pass filter that dampens high-curvature noise to reveal the optimal descent direction. Capitalizing on this insight, we propose Extra-Merge, a training-free strategy that extrapolates along this subspace to minimize loss without additional gradient updates. Extensive experiments across GPT-2 and LLaMA families (124M to 2B) demonstrate that Extra-Merge consistently outperforms standard merging baselines. Notably, it yields consistent zero-shot accuracy gains on Pythia-12B downstream tasks and generalizes effectively to the Muon optimizer (Jordan et al., 2024).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10a3b1c4-7d24-400e-a15e-70713237be0cBuilds on20
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
Related papers
- Revisiting the Role of Pretrained Weights in Model Merging: On Near-Optimality within the Core SubspaceWenju Sun, Qingyong Li, Tiancheng Li, Yangliao Geng et al.ICML 2026
- Unraveling LoRA Interference: Orthogonal Subspaces for Robust Model MergingHaobo Zhang, Jiayu ZhouACL 2025
- Maximizing Intermediate Checkpoint Value in LLM Pretraining with Bayesian OptimizationDeyuan Liu, Zecheng Wang, Bingning Wang, Weipeng Chen et al.ICML 2025
- Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer ChunkingDengming Zhang, Xiaowen Ma, Zhenliang Ni, Zhenkai Wu et al.ICLR 2026 · 6 citations
- CAMEx: Curvature-aware Merging of ExpertsViet Dung Nguyen, Minh Nguyen Hoang, Luc Q. Nguyen, Rachel S. Y. Teo et al.ICLR 2025
