Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
Jeonghoon Kim, Byeongchan Lee, Cheonbok Park, Yeontaek Oh, Beomjun Kim, Taehwan Yoo, Seongjin Shin, Dongyoon Han, Jinwoo Shin, Kang Min Yoo
摘要
Selecting a layer normalization (LN) strategy that stabilizes training and speeds convergence in Transformers remains difficult, even for today’s large language models (LLM). We present a comprehensive analytical foundation for understanding how different LN strategies influence training dynamics in large-scale Transformers. Until recently, Pre-LN and Post-LN have long dominated practices despite their limitations in large-scale training. However, several open-source models have recently begun silently adopting a third strategy without much explanation. This strategy places normalization layer peripherally around sublayers, a design we term Peri-LN. While Peri-LN has demonstrated promising performance, its precise mechanisms and benefits remain almost unexplored. Our in-depth analysis delineates the distinct behaviors of LN strategies, showing how each placement shapes activation variance and gradient propagation. To validate our theoretical insight, we conduct extensive experiments on Transformers up to B parameters, showing that Peri-LN consistently achieves more balanced variance growth, steadier gradient flow, and convergence stability. Our results suggest that Peri-LN warrants broader consideration for large-scale Transformer architectures, providing renewed insights into the optimal placement of LN.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SiameseNorm: Breaking the Barrier to Reconciling Pre/Post-NormTianyu Li, Dongchen Han, Zixuan Cao, Haofeng Huang 等ICML 2026 · 被引用 7 次
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information LossBozhou Li, Xinda Xue, Sihan Yang, Yang Shi 等ICLR 2026 · 被引用 5 次
- GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation ScalingTianhao Chen, Xin Xu, Zijing Liu, Pengxiang Li 等NeurIPS 2025 · 被引用 2 次
- UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose EstimationHaopeng Chen, Yihao Ai, Kabeen Kim, Robby T. Tan 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
相关 Paper
- HybridNorm: Towards Stable and Efficient Transformer Training via Hybrid NormalizationZhijian Zhuo, Yutao Zeng, Ya Wang, Sijun Zhang 等NeurIPS 2025 · 被引用 23 次
- Normalization in Attention DynamicsNikita Karagodin, Shu Ge, Yury Polyanskiy, Philippe RigolletNeurIPS 2025 · 被引用 10 次
- SpanNorm: Reconciling Training Stability and Performance in Deep TransformersChao Wang, Bei Li, Jiaqi Zhang, Xinyu Liu 等ICML 2026
- Pre-RMSNorm and Pre-CRMSNorm Transformers: Equivalent and Efficient Pre-LN TransformersZixuan Jiang, Jiaqi Gu, Hanqing Zhu, David Z. PanNeurIPS 2023 · 被引用 44 次
- Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LNPengxiang Li, Lu Yin, Shiwei LiuICLR 2025
