Lune

ICML2026Top-tier venue

Controlled LLM Training on Spectral Sphere

Tian Xie, Haoming Luo, Haoyu Tang, Hu Yiwen, Jason Liu, Qingnan Ren, Yang Wang, Xin Zhao, Rui Yan, Bing Su, Chong Luo, Baining Guo

2026Year
22Citations
4Top-tier citations

Abstract

Scaling large models requires optimization strategies that ensure rapid convergence grounded in stability. Maximal Update Parametrization (μ\boldsymbol{\mu}P) provides a theoretical safeguard for width-invariant Θ(1)\Theta(1) activation control, whereas emerging optimizers like Muon are only "half-aligned" with these constraints: they control updates but allow weights to drift. To address this limitation, we introduce the Spectral Sphere Optimizer (SSO), which enforces strict module-wise spectral constraints on both weights and their updates. By deriving the steepest descent direction on the spectral sphere, SSO realizes a fully μ\boldsymbol{\mu}P-aligned optimization process. To enable large‑scale training, we implement SSO as an efficient parallel algorithm within Megatron. Through extensive pretraining on diverse architectures, including Dense 1.7B, MoE 8B-A1B, and 200-layer DeepNet models, SSO consistently outperforms AdamW and Muon. Furthermore, we observe significant practical stability benefits, including improved MoE router load balancing, suppressed outliers, and strictly bounded activations.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 953636d9-0db6-4e38-bc74-0d143862be96

Cited by top-tier papers4

Ask how each one uses it

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines