Tensor Programs IIb: Architectural Universality Of Neural Tangent Kernel Training Dynamics
Greg Yang, Etai Littwin
摘要
Yang (2020a) recently showed that the Neural Tangent Kernel (NTK) at initialization has an infinite-width limit for a large class of architectures including modern staples such as ResNet and Transformers. However, their analysis does not apply to training. Here, we show the same neural networks (in the so-called NTK parametrization) during training follow a kernel gradient descent dynamics in function space, where the kernel is the infinite-width NTK. This completes the proof of the architectural universality of NTK behavior. To achieve this result, we apply the Tensor Programs technique: Write the entire SGD dynamics inside a Tensor Program and analyze it via the Master Theorem. To facilitate this proof, we develop a graphical notation for Tensor Programs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc 等ICML 2023 · 被引用 260 次
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 被引用 140 次
- A Kernel-Based View of Language Model Fine-TuningSadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen 等ICML 2023 · 被引用 111 次
- DAVINZ: Data Valuation using Deep Neural Networks at InitializationZhaoxuan Wu, Yao Shu, Bryan Kian Hsiang LowICML 2022 · 被引用 71 次
它引用的顶会 Paper3
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 被引用 169 次
- Infinite attention: NNGP and NTK for deep attention networksJiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, Roman NovakICML 2020 · 被引用 147 次
- The Recurrent Neural Tangent KernelSina Alemohammad, Zichao Wang, Randall Balestriero, Richard G. BaraniukICLR 2021 · 被引用 6 次
相关 Paper
- Adaptive Optimization in the ∞-Width LimitEtai Littwin, Greg YangICLR 2023
- Real-Valued Backpropagation is Unsuitable for Complex-Valued Neural NetworksZhi-Hao Tan, Yi Xie, Yuan Jiang, Zhi-Hua ZhouNeurIPS 2022 · 被引用 16 次
- Global Convergence and Rich Feature Learning in L-Layer Infinite-Width Neural Networks under μ ParametrizationZixiang Chen, Greg Yang, Qingyue Zhao, Quanquan GuICML 2025
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 被引用 167 次
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
