Extrapolation and Spectral Bias of Neural Nets with Hadamard Product: a Polynomial Net Study
Yongtao Wu, Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan Cevher
Abstract
Neural tangent kernel (NTK) is a powerful tool to analyze training dynamics of neural networks and their generalization bounds. The study on NTK has been devoted to typical neural network architectures, but it is incomplete for neural networks with Hadamard products (NNs-Hp), e.g., StyleGAN and polynomial neural networks (PNNs). In this work, we derive the finite-width NTK formulation for a special class of NNs-Hp, i.e., polynomial neural networks. We prove their equivalence to the kernel regression predictor with the associated NTK, which expands the application scope of NTK. Based on our results, we elucidate the separation of PNNs over standard neural networks with respect to extrapolation and spectral bias. Our two key insights are that when compared to standard neural networks, PNNs can fit more complicated functions in the extrapolation regime and admit a slower eigenvalue decay of the respective NTK, leading to a faster learning towards high-frequency functions. Besides, our theoretical results can be extended to other types of NNs-Hp, which expand the scope of our work. Our empirical results validate the separations in broader classes of NNs-Hp, which provide a good justification for a deeper understanding of neural architectures. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a3576a8-795b-419c-bb1b-7b3f304c20afCited by top-tier papers3
- TMPNN: High-Order Polynomial Regression Based on Taylor Map FactorizationAndrei Ivanov, Stefan Maria AiluroAAAI 2024 · 3 citations
- Regularization of polynomial networks for image recognitionGrigorios G. Chrysos, Bohan Wang, Jiankang Deng, Volkan CevherCVPR 2023
- Deep Neural Networks Tend To Extrapolate PredictablyKatie Kang, Amrith Setlur, Claire J. Tomlin, Sergey LevineICLR 2024
Builds on14
- How Neural Networks Extrapolate: From Feedforward to Graph Neural NetworksKeyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du et al.ICLR 2021 · 364 citations
- The Risks of Invariant Risk MinimizationElan Rosenfeld, Pradeep Kumar Ravikumar, Andrej RisteskiICLR 2021 · 356 citations
- Multiplicative Filter NetworksRizal Fathony, Anit Kumar Sahu, Devin Willmott, J. Zico KolterICLR 2021 · 185 citations
- Multiplicative Interactions and Where to Find ThemSiddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz et al.ICLR 2020 · 152 citations
- Convolutional Tensor-Train LSTM for Spatio-Temporal LearningJiahao Su, Wonmin Byeon, Jean Kossaifi, Furong Huang et al.NeurIPS 2020 · 146 citations
Related papers
- The Spectral Bias of Polynomial Neural NetworksMoulik Choraria, Leello Tadesse Dadi, Grigorios Chrysos, Julien Mairal et al.ICLR 2022 · 26 citations
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 245 citations
- Separable Neural Networks: Approximation Theory, NTK Regime, and Preconditioned Gradient DescentYisi Luo, Deyu MengICLR 2026
- Finite-Width Neural Tangent Kernels from Feynman DiagramsMax Guillen, Philipp Misof, Jan GerkenICML 2026 · 1 citation
