Lune

ICML2026Top-tier venue

Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training

WenJie Zhou, Bohan Wang, Hongtao Zhang, Chenxi Jia, Wei Chen, Xueqi Cheng

2026Year

Abstract

Model merging has emerged as a lightweight paradigm for enhancing Large Language Models (LLMs), yet its underlying mechanisms remain poorly understood. In this work, we analyze late-stage pre-training trajectories and uncover a Rank-1 Subspace phenomenon: while raw optimization steps oscillate violently, consecutive merged checkpoints collapse onto a stable, approximately one-dimensional linear manifold. We theoretically ground this observation in a river-valley landscape analysis: averaging acts as a geometric low-pass filter that dampens high-curvature noise to reveal the optimal descent direction. Capitalizing on this insight, we propose Extra-Merge, a training-free strategy that extrapolates along this subspace to minimize loss without additional gradient updates. Extensive experiments across GPT-2 and LLaMA families (124M to 2B) demonstrate that Extra-Merge consistently outperforms standard merging baselines. Notably, it yields consistent zero-shot accuracy gains on Pythia-12B downstream tasks and generalizes effectively to the Muon optimizer (Jordan et al., 2024).

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 10a3b1c4-7d24-400e-a15e-70713237be0c

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines