Lune

ICML2026Top-tier venue

Escaping the Subspace Trap: The Role of Optimizer Geometry in Model Width Expansion

Jiabei Chen, Haoyu Wang, Yang Yu, Yao Xu, Liangdong Wang, Guang Liu, Shizhu He, Jun Zhao, Kang Liu

2026Year

Abstract

Pre-training large language models from scratch is prohibitively expensive as model scales increase. A practical alternative is Model Width Expansion (MWE), which grows a larger model from a well-pretrained ''seed'' model to inherit existing capabilities at initialization. However, we identify a phenomenon termed the Subspace Trap : during continual pre-training, parameter updates largely stagnate within a low-dimensional subspace aligned with the initialization, limiting the effective capacity of the expanded model. Our theoretical analysis investigates this issue by attributing it to the function-preserving properties of width expansion. In particular, element-wise adaptive optimizers remain confined to the trap, whereas optimizers that yield an isotropic geometry of parameter updates can escape. To demonstrate the impact of the subspace trap on model performance, we conduct empirical experiments across different model sizes and model families, which show that escaping the trap is principally effective in improving training efficiency and overall model performance. Detailed mechanistic analyses further confirm that escaping the trap indeed activates the new dimensions to encode general knowledge. Our code is available at https://github.com/A-PolarBear/Model-Width-Expansion.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 039e3c49-1f3c-4348-a4da-573cc05079f8

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines