Separable Neural Networks: Approximation Theory, NTK Regime, and Preconditioned Gradient Descent
Yisi Luo, Deyu Meng
Abstract
Separable neural networks (SepNNs) are emerging neural architectures that significantly reduce computational costs by factorizing a multivariate function into linear combinations of univariate functions, benefiting downstream applications such as implicit neural representations (INRs) and physics-informed neural networks (PINNs). However, fundamental theoretical analysis for SepNN, including detailed representation capacity and spectral bias characterization & alleviation, remains unexplored. This work makes three key contributions to theoretically understanding and improving SepNN. First, using Weierstrass-based approximation and universal approximation theory, we prove that SepNN can approximate any multivariate function with arbitrary precision, confirming its representation completeness. Second, we derive the neural tangent kernel (NTK) regimes for SepNN, showing that the NTK of infinite-width SepNN converges to a deterministic (or random) kernel under infinite (or fixed) decomposition rank, with corresponding convergence and spectral bias characterization. Third, we propose an efficient separable preconditioned gradient descent (SepPGD) for optimizing SepNN, which alleviates the spectral bias of SepNN by provably adjusting its NTK spectrum. The SepPGD enjoys an efficient complexity for training samples, which is much more efficient than previous neural network PGD methods. Extensive experiments for kernel ridge regression, image and surface representation using INRs, and numerical PDEs using PINNs validate the efficiency of SepNN and the effectiveness of SepPGD for alleviating spectral bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f82d7d9-fe47-4d24-b806-316ea255a8baBuilds on13
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 183 citations
- Separable Physics-Informed Neural NetworksJunwoo Cho, Seungtae Nam, Hyunmo Yang, Seok-Bae Yun et al.NeurIPS 2023 · 138 citations
- Fast Finite Width Neural Tangent KernelRoman Novak, Jascha Sohl-Dickstein, Samuel S. SchoenholzICML 2022 · 72 citations
Related papers
- Deep Learning with Learnable Product-Structured ActivationsSaanjali Maharaj, Prasanth B. NairICLR 2026
- Extrapolation and Spectral Bias of Neural Nets with Hadamard Product: a Polynomial Net StudyYongtao Wu, Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos et al.NeurIPS 2022 · 19 citations
- The Challenges of the Nonlinear Regime for Physics-Informed Neural NetworksAndrea Bonfanti, Giuseppe Bruno, Cristina CiprianiNeurIPS 2024 · 41 citations
- How Learnable Grids Recover Fine Detail in Low Dimensions: A Neural Tangent Kernel Analysis of Multigrid Parametric EncodingsSamuel Audia, Soheil Feizi, Matthias Zwicker, Dinesh ManochaICLR 2025
- The Spectral Bias of Polynomial Neural NetworksMoulik Choraria, Leello Tadesse Dadi, Grigorios Chrysos, Julien Mairal et al.ICLR 2022 · 26 citations
