Non-Gaussian Tensor Programs
Eugene A. Golikov, Greg Yang
摘要
Does it matter whether one randomly initializes a neural network (NN) from Gaussian, uniform, or other distributions? We show the answer is ”yes” in some parameter tensors (the so-called matrix-like parameters) but ”no” in others when the NN is wide. This is a specific instance of a more general universality principle for Tensor Programs (TP) that informs precisely when the limit of a program depends on the distribution of its initial matrices and vectors. To obtain this principle, we develop the theory of non-Gaussian Tensor Programs. As corollaries, we obtain all previous consequences of the TP framework (such as NNGP/NTK correspondence, Free Indepedence Principle, Dynamical Dichotomy Theorem, and µ -parametrization) for NNs with non-Gaussian weights. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Principled Weight Initialisation for Input-Convex Neural NetworksPieter-Jan Hoedt, Günter KlambauerNeurIPS 2023 · 被引用 19 次
- Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor ProgramsMana Sakai, Ryo Karakida, Masaaki ImaizumiNeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
- Tensor Programs IIb: Architectural Universality Of Neural Tangent Kernel Training DynamicsGreg Yang, Etai LittwinICML 2021 · 被引用 81 次
- Towards a General Theory of Infinite-Width Limits of Neural ClassifiersEugene A. GolikovICML 2020 · 被引用 10 次
相关 Paper
- On the Impacts of the Random Initialization in the Neural Tangent Kernel TheoryGuhan Chen, Yicheng Li, Qian LinNeurIPS 2024 · 被引用 7 次
- Beyond IID weights: sparse and low-rank deep Neural Networks are also Gaussian ProcessesThiziri Nait Saada, Alireza Naderi, Jared TannerICLR 2024 · 被引用 2 次
- Global Convergence and Rich Feature Learning in L-Layer Infinite-Width Neural Networks under μ ParametrizationZixiang Chen, Greg Yang, Qingyue Zhao, Quanquan GuICML 2025
- A Unified Weight Initialization Paradigm for Tensorial Convolutional Neural NetworksYu Pan, Zeyong Su, Ao Liu, Jingquan Wang 等ICML 2022 · 被引用 15 次
- Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning SystemsOri Shem-Ur, Khen Cohen, Yaron OzICLR 2026
