Conservation Laws for Modern Neural Architectures
Viet Hoang Tran, VINH KHANH BUI, Ngoc Tan Lai, Nam Nguyen, Tuan Dam, Tan Nguyen
摘要
Understanding gradient descent dynamics is key to explaining the success of over-parameterized models, where implicit bias manifests through conservation laws in gradient flow. While such laws are well understood for linear and ReLU networks, they remain largely unexplored for modern architectures. This work develops a unified framework to characterize conservation laws for contemporary models, including feedforward networks with GELU, SiLU, and SwiGLU activations, multihead attention with sinusoidal and rotary positional encodings, and Mixture-of-Experts architectures under diverse gating designs. Our theoretical findings are supported by experiments that validate the predicted invariants.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BASE Layers: Simplifying Training of Large, Sparse ModelsMike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal 等ICML 2021 · 被引用 382 次
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu 等ICLR 2023 · 被引用 234 次
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning DynamicsDaniel Kunin, Javier Sagastuy-Breña, Surya Ganguli, Daniel L. K. Yamins 等ICLR 2021 · 被引用 100 次
- Understanding the Dynamics of Gradient Flow in Overparameterized Linear modelsSalma Tarmoun, Guilherme França, Benjamin D. Haeffele, René VidalICML 2021 · 被引用 76 次
相关 Paper
- Transformative or Conservative? Conservation laws for ResNets and TransformersSibylle Marcotte, Rémi Gribonval, Gabriel PeyréICML 2025
- Abide by the law and follow the flow: conservation laws for gradient flowsSibylle Marcotte, Rémi Gribonval, Gabriel PeyréNeurIPS 2023 · 被引用 54 次
- Keep the Momentum: Conservation Laws beyond Euclidean Gradient FlowsSibylle Marcotte, Rémi Gribonval, Gabriel PeyréICML 2024 · 被引用 8 次
- Intrinsic training dynamics of deep neural networksSibylle Marcotte, Gabriel Peyré, Rémi GribonvalICLR 2026 · 被引用 4 次
- Avoiding Kernel Fixed Points: Computing with ELU and GELU Infinite NetworksRussell Tsuchida, Tim Pearce, Christopher van der Heide, Fred Roosta 等AAAI 2021 · 被引用 10 次
