Fast Convergence of Natural Gradient Descent for Over-parameterized Physics-Informed Neural Networks
Xianliang Xu, Wang Kong, Jiaheng Mao, Zhongyi Huang, Ye Li
Abstract
In the context of over-parameterization, there is a line of work demonstrating that randomly initialized (stochastic) gradient descent (GD) converges to a globally optimal solution at a linear convergence rate for the quadratic loss function. However, the convergence rate of GD for training two-layer neural networks exhibits poor dependence on the sample size and the Gram matrix, leading to a slow training process. In this paper, we show that for training two-layer Physics-Informed Neural Networks (PINNs), the learning rate can be improved from the smallest eigenvalue of the limiting Gram matrix to the reciprocal of the largest eigenvalue, implying that GD actually enjoys a faster convergence rate. Despite such improvements, the convergence rate is still tied to the least eigenvalue of the Gram matrix, leading to slow convergence. We then develop the positive definiteness of Gram matrices with general smooth activation functions and provide the convergence analysis of natural gradient descent (NGD) in training two-layer PINNs, demonstrating that the maximal learning rate can be and at this rate, the convergence rate is independent of the Gram matrix. In particular, for smooth activation functions, the convergence rate of NGD is quadratic. Numerical experiments are conducted to verify our theoretical results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7db47131-a472-4e2e-bec0-02ecad9fe2b9Builds on5
- Challenges in Training PINNs: A Loss Landscape PerspectivePratik Rathore, Weimu Lei, Zachary Frangella, Lu Lu et al.ICML 2024 · 137 citations
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU NetworksQuynh Nguyen, Marco Mondelli, Guido F. MontúfarICML 2021 · 98 citations
- How does PDE order affect the convergence of PINNs?Changhoon Song, Yesom Park, Myungjoo KangNeurIPS 2024 · 17 citations
- Achieving High Accuracy with PINNs via Energy Natural Gradient DescentJohannes Müller, Marius ZeinhoferICML 2023 · 13 citations
- Gradient Descent Finds the Global Optima of Two-Layer Physics-Informed Neural NetworksYihang Gao, Yiqi Gu, Michael NgICML 2023 · 12 citations
Related papers
- Near-optimal Sketchy Natural Gradients for Physics-Informed Neural NetworksMaricela Best McKay, Avleen Kaur, Chen Greif, Brian WettonICML 2025
- Implicit Stochastic Gradient Descent for Training Physics-Informed Neural NetworksYe Li, Songcan Chen, Sheng-Jun HuangAAAI 2023 · 5 citations
- An operator preconditioning perspective on training in physics-informed machine learningTim De Ryck, Florent Bonnet, Siddhartha Mishra, Emmanuel de BézenacICLR 2024 · 28 citations
- Improving Energy Natural Gradient Descent through Woodbury, Momentum, and RandomizationAndrés Guzmán-Cordero, Felix Dangel, Gil Goldshlager, Marius ZeinhoferNeurIPS 2025 · 17 citations
- The Challenges of the Nonlinear Regime for Physics-Informed Neural NetworksAndrea Bonfanti, Giuseppe Bruno, Cristina CiprianiNeurIPS 2024 · 41 citations
