Lune

ICML2026Top-tier venue

Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training

Seyed Morteza Emadi

2026Year
1Citations

Abstract

Attention scores in transformers are bilinear forms Sij=xi⊤Mxj/dhS_{ij} = x_i^\top M x_j / \sqrt{d_h} whose maximum magnitude governs overflow risk in low-precision training. We derive a rank-aware concentration inequality: when the interaction matrix M=WQWK⊤M = W^Q W^{K\top} has rank r≪dr \ll d, tail probabilities for max⁡i,j∣Sij∣\max_{i,j}|S_{ij}| decay as exp⁡(−d2α2/(γr))\exp(-d^{2}\alpha^{2}/(\gamma r)) rather than exp⁡(−dα2)\exp(-d\alpha^{2}), where γ>1\gamma > 1. For transformer attention where r=dhr = d_h, this yields 88--28×28\times tighter concentration than rank-agnostic bounds in modern architectures. We apply this result to FP8 training, deriving geometry-aware scale factors that provide principled overflow guarantees without observing activations. The method computes per-layer scales from the spectral norm ∥WQWK⊤∥2\|W^Q W^{K\top}\|_2 via implicit power iteration, includes a grouped query attention formulation that avoids key expansion, and remains compatible with fused attention kernels. Across GPT-2 XL to Llama-2-70B, geometry-aware scaling eliminates overflows in transient scenarios where delayed scaling fails, while achieving comparable downstream MMLU accuracy.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 245cbe0b-82a9-4f7a-b382-bf5c7c634f6e

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines