Lune

ICML2026顶会

Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training

Seyed Morteza Emadi

2026年份
1被引次数

摘要

Attention scores in transformers are bilinear forms Sij=xi⊤Mxj/dhS_{ij} = x_i^\top M x_j / \sqrt{d_h} whose maximum magnitude governs overflow risk in low-precision training. We derive a rank-aware concentration inequality: when the interaction matrix M=WQWK⊤M = W^Q W^{K\top} has rank r≪dr \ll d, tail probabilities for max⁡i,j∣Sij∣\max_{i,j}|S_{ij}| decay as exp⁡(−d2α2/(γr))\exp(-d^{2}\alpha^{2}/(\gamma r)) rather than exp⁡(−dα2)\exp(-d\alpha^{2}), where γ>1\gamma > 1. For transformer attention where r=dhr = d_h, this yields 88--28×28\times tighter concentration than rank-agnostic bounds in modern architectures. We apply this result to FP8 training, deriving geometry-aware scale factors that provide principled overflow guarantees without observing activations. The method computes per-layer scales from the spectral norm ∥WQWK⊤∥2\|W^Q W^{K\top}\|_2 via implicit power iteration, includes a grouped query attention formulation that avoids key expansion, and remains compatible with fused attention kernels. Across GPT-2 XL to Llama-2-70B, geometry-aware scaling eliminates overflows in transient scenarios where delayed scaling fails, while achieving comparable downstream MMLU accuracy.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper7

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖