Lune

ICML2026顶会

Controlled LLM Training on Spectral Sphere

Tian Xie, Haoming Luo, Haoyu Tang, Hu Yiwen, Jason Liu, Qingnan Ren, Yang Wang, Xin Zhao, Rui Yan, Bing Su, Chong Luo, Baining Guo

2026年份
22被引次数
4顶会引用

摘要

Scaling large models requires optimization strategies that ensure rapid convergence grounded in stability. Maximal Update Parametrization (μ\boldsymbol{\mu}P) provides a theoretical safeguard for width-invariant Θ(1)\Theta(1) activation control, whereas emerging optimizers like Muon are only "half-aligned" with these constraints: they control updates but allow weights to drift. To address this limitation, we introduce the Spectral Sphere Optimizer (SSO), which enforces strict module-wise spectral constraints on both weights and their updates. By deriving the steepest descent direction on the spectral sphere, SSO realizes a fully μ\boldsymbol{\mu}P-aligned optimization process. To enable large‑scale training, we implement SSO as an efficient parallel algorithm within Megatron. Through extensive pretraining on diverse architectures, including Dense 1.7B, MoE 8B-A1B, and 200-layer DeepNet models, SSO consistently outperforms AdamW and Muon. Furthermore, we observe significant practical stability benefits, including improved MoE router load balancing, suppressed outliers, and strictly bounded activations.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper4

问问它们各自怎么用它

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖