Non-Gaussian Tensor Programs
Eugene A. Golikov, Greg Yang
Abstract
Does it matter whether one randomly initializes a neural network (NN) from Gaussian, uniform, or other distributions? We show the answer is ”yes” in some parameter tensors (the so-called matrix-like parameters) but ”no” in others when the NN is wide. This is a specific instance of a more general universality principle for Tensor Programs (TP) that informs precisely when the limit of a program depends on the distribution of its initial matrices and vectors. To obtain this principle, we develop the theory of non-Gaussian Tensor Programs. As corollaries, we obtain all previous consequences of the TP framework (such as NNGP/NTK correspondence, Free Indepedence Principle, Dynamical Dichotomy Theorem, and µ -parametrization) for NNs with non-Gaussian weights. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2970ca4-a746-4a7e-91db-9a58c760815bCited by top-tier papers2
- Principled Weight Initialisation for Input-Convex Neural NetworksPieter-Jan Hoedt, Günter KlambauerNeurIPS 2023 · 19 citations
- Infinite-Width Limit of a Single Attention Layer: Analysis via Tensor ProgramsMana Sakai, Ryo Karakida, Masaaki ImaizumiNeurIPS 2025 · 5 citations
Builds on4
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Tensor Programs IIb: Architectural Universality Of Neural Tangent Kernel Training DynamicsGreg Yang, Etai LittwinICML 2021 · 81 citations
- Towards a General Theory of Infinite-Width Limits of Neural ClassifiersEugene A. GolikovICML 2020 · 10 citations
Related papers
- On the Impacts of the Random Initialization in the Neural Tangent Kernel TheoryGuhan Chen, Yicheng Li, Qian LinNeurIPS 2024 · 7 citations
- Beyond IID weights: sparse and low-rank deep Neural Networks are also Gaussian ProcessesThiziri Nait Saada, Alireza Naderi, Jared TannerICLR 2024 · 2 citations
- Global Convergence and Rich Feature Learning in L-Layer Infinite-Width Neural Networks under μ ParametrizationZixiang Chen, Greg Yang, Qingyue Zhao, Quanquan GuICML 2025
- A Unified Weight Initialization Paradigm for Tensorial Convolutional Neural NetworksYu Pan, Zeyong Su, Ao Liu, Jingquan Wang et al.ICML 2022 · 15 citations
- Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning SystemsOri Shem-Ur, Khen Cohen, Yaron OzICLR 2026
