Lune

ICML2026顶会

GradPower: Powering Gradients for Faster Language Model Pre-Training

Jinbo Wang, Mingze Wang, Jiaqi Zhang, Wei Wang, Peng Pei, Xunliang Cai, Weinan E, Lei Wu

2026年份
4被引次数

摘要

We propose GradPower, a lightweight gradient-transformation technique for accelerating language model pre-training. Given a gradient vector g=(gi)i\boldsymbol{g}=(g_ {i})_ {i}, GradPower first applies the elementwise sign-power transformation: φp(g)=(sign(gi)∣gi∣p)i\varphi_ p(\boldsymbol{g}) = \left({\rm sign}(g_ i)|g_ i|^p\right)_ {i} for a fixed p>0p>0, and then feeds the transformed gradient into a base optimizer. Notably, GradPower requires only a single-line code change and no modifications to the base optimizer’s internal logic, including the hyperparameters. When applied to AdamW (termed AdamWPower), GradPower consistently achieves lower terminal loss across diverse architectures (LLaMA, Qwen2MoE), parameter scales (66M to 2B), datasets (C4, OpenWebText), and learning-rate schedules (cosine, warmup-stable-decay). The most pronounced gains are observed when training modern mixture-of-experts models with warmup-stable-decay schedules. GradPower also integrates seamlessly with other state-of-the-art optimizers, such as Muon, yielding further improvements. Finally, we provide theoretical analyses that reveal the underlying mechanism of GradPower and highlight the influence of gradient noise.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 2febfac7-44df-436f-b877-9498070fdff7

它引用的顶会 Paper23

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖