Transformer Block Coupling and its Correlation with Generalization in LLMs
Murdock Aubry, Haoming Meng, Anton Sugolov, Vardan Papyan
摘要
Large Language Models (LLMs) have made significant strides in natural language processing, and a precise understanding of the internal mechanisms driving their success is essential. In this work, we trace the trajectories of individual tokens as they pass through transformer blocks, and linearize the system along these trajectories through their Jacobian matrices. By examining the relationships between these Jacobians, we uncover a phenomenon in a variety of LLMs, characterized by the coupling of their top singular vectors across tokens and depth. Our findings reveal that coupling with model performance, and that this relationship is stronger than with other hyperparameters, namely parameter budget, model depth, and embedding dimension. We further investigate the emergence of these properties through training, noting the development of coupling, as well as an increase in linearity and layer-wise exponential growth in the token trajectories. These collective insights provide a novel perspective on the interactions between token embeddings, and prompt further approaches to study training and generalization in LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Short-LVLM: Compressing and Accelerating Large Vision-Language Models by Pruning Redundant LayersJi Ma, Wei Suo, Peng Wang, Yanning ZhangACM MM 2025 · 被引用 9 次
- Mining Useful General Data for Low-Resource Domain AdaptationPingjie Wang, Hongcheng Liu, Yusheng Liao, Ziqing Fan 等ICML 2026 · 被引用 3 次
- Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal ControlJulian Skifstad, Xinyue Annie Yang, Glen ChouICML 2026 · 被引用 2 次
- Gradient Smoothing: Coupling Layer-wise Updates for Improved OptimizationHaoming Meng, Anton Sugolov, Vardan PapyanICML 2026
它引用的顶会 Paper26
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
相关 Paper
- The Role of Sparsity for Length Generalization in LLMsNoah Golowich, Samy Jelassi, David Brandfonbrener, Sham M. Kakade 等ICML 2025
- Lines of Thought in Large Language ModelsRaphaël Sarfati, Toni J. B. Liu, Nicolas Boullé, Christopher J. EarlsICLR 2025
- Over-Tokenized Transformer: Vocabulary is Generally Worth ScalingHongzhi Huang, Defa Zhu, Banggu Wu, Yutao Zeng 等ICML 2025
- Incorporating Residual and Normalization Layers into Analysis of Masked Language ModelsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2021 · 被引用 28 次
- Demystifying Singular Defects in Large Language ModelsHaoqi Wang, Tong Zhang, Mathieu SalzmannICML 2025
