Spectral Scaling Laws in Language Models: emphHow Effectively Do Feed-Forward Networks Use Their Latent Space?
Nandan Kumar Jha, Brandon Reagen
摘要
As Large Language Models (LLMs) scale, the question is not just how large they become, but how much of their capacity is effectively utilized. Existing scaling laws relate model size to loss, yet overlook how components exploit their latent space. In this work, we focus on Feed-Forward Networks (FFNs) and recast width selection as a spectral utilization optimization problem. Using a lightweight diagnostic suite: Hard Rank (participation ratio), Soft Rank (Shannon Rank), Spectral Concentration, and the composite Spectral Utilization Index (SUI), we quantify how many latent directions are meaningfully activated across LLaMA, GPT-2, and nGPT families. Our key finding is an Asymmetric Spectral Scaling Law: soft rank follows an almost perfect power law with FFN width, while hard rank grows only sublinearly, with high variance. This asymmetry suggests that widening FFNs mostly adds low-energy tail directions, while dominant-mode subspaces saturate early. Moreover, at larger widths, variance further collapses into a narrow subspace, leaving much of the latent space under-utilized. These results recast FFN width selection as a principled trade-off between tail capacity and dominant-mode capacity, offering concrete guidance for inference-efficient LLM design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper21
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling LawsNikhil Sardana, Jacob P. Portes, Sasha Doubov, Jonathan FrankleICML 2024 · 被引用 144 次
- Scaling Laws with Vocabulary: Larger Models Deserve Larger VocabulariesChaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff 等NeurIPS 2024 · 被引用 135 次
- RankMe: Assessing the Downstream Performance of Pretrained Self-Supervised Representations by Their RankQuentin Garrido, Randall Balestriero, Laurent Najman, Yann LeCunICML 2023 · 被引用 127 次
- 4+3 Phases of Compute-Optimal Neural Scaling LawsElliot Paquette, Courtney Paquette, Lechao Xiao, Jeffrey PenningtonNeurIPS 2024 · 被引用 70 次
相关 Paper
- NerVE: Nonlinear Eigenspectrum Dynamics in LLM Feed-Forward NetworksNandan Kumar Jha, Brandon ReagenICLR 2026 · 被引用 4 次
- Spectra: Rethinking Optimizers for LLMs Under Spectral AnisotropyZhendong Huang, Hengjie Cao, Fang DONG(董方), Ruijun Huang 等ICML 2026 · 被引用 5 次
- Eigenspectrum Analysis of Neural Networks without Aspect Ratio BiasYuanzhe Hu, Kinshuk Goel, Vlad Killiakov, Yaoqing YangICML 2025
- Spectral Signatures of Large Language ModelsZhuoying Zhang, Ishan V. Prasad, Yuanzhe Hu, Zihang Liu 等KDD 2026
- Spectral Reach: Understanding Neural Scaling as Progress into the Spectral TailKonstantin Nikolaou, Jonas Scheunemann, Sven Krippendorf, Samuel Tovey 等ICML 2026
