Fast Convergence of Natural Gradient Descent for Over-parameterized Physics-Informed Neural Networks
Xianliang Xu, Wang Kong, Jiaheng Mao, Zhongyi Huang, Ye Li
摘要
In the context of over-parameterization, there is a line of work demonstrating that randomly initialized (stochastic) gradient descent (GD) converges to a globally optimal solution at a linear convergence rate for the quadratic loss function. However, the convergence rate of GD for training two-layer neural networks exhibits poor dependence on the sample size and the Gram matrix, leading to a slow training process. In this paper, we show that for training two-layer Physics-Informed Neural Networks (PINNs), the learning rate can be improved from the smallest eigenvalue of the limiting Gram matrix to the reciprocal of the largest eigenvalue, implying that GD actually enjoys a faster convergence rate. Despite such improvements, the convergence rate is still tied to the least eigenvalue of the Gram matrix, leading to slow convergence. We then develop the positive definiteness of Gram matrices with general smooth activation functions and provide the convergence analysis of natural gradient descent (NGD) in training two-layer PINNs, demonstrating that the maximal learning rate can be and at this rate, the convergence rate is independent of the Gram matrix. In particular, for smooth activation functions, the convergence rate of NGD is quadratic. Numerical experiments are conducted to verify our theoretical results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- Challenges in Training PINNs: A Loss Landscape PerspectivePratik Rathore, Weimu Lei, Zachary Frangella, Lu Lu 等ICML 2024 · 被引用 137 次
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU NetworksQuynh Nguyen, Marco Mondelli, Guido F. MontúfarICML 2021 · 被引用 98 次
- How does PDE order affect the convergence of PINNs?Changhoon Song, Yesom Park, Myungjoo KangNeurIPS 2024 · 被引用 17 次
- Achieving High Accuracy with PINNs via Energy Natural Gradient DescentJohannes Müller, Marius ZeinhoferICML 2023 · 被引用 13 次
- Gradient Descent Finds the Global Optima of Two-Layer Physics-Informed Neural NetworksYihang Gao, Yiqi Gu, Michael NgICML 2023 · 被引用 12 次
相关 Paper
- Near-optimal Sketchy Natural Gradients for Physics-Informed Neural NetworksMaricela Best McKay, Avleen Kaur, Chen Greif, Brian WettonICML 2025
- Implicit Stochastic Gradient Descent for Training Physics-Informed Neural NetworksYe Li, Songcan Chen, Sheng-Jun HuangAAAI 2023 · 被引用 5 次
- An operator preconditioning perspective on training in physics-informed machine learningTim De Ryck, Florent Bonnet, Siddhartha Mishra, Emmanuel de BézenacICLR 2024 · 被引用 28 次
- Improving Energy Natural Gradient Descent through Woodbury, Momentum, and RandomizationAndrés Guzmán-Cordero, Felix Dangel, Gil Goldshlager, Marius ZeinhoferNeurIPS 2025 · 被引用 17 次
- The Challenges of the Nonlinear Regime for Physics-Informed Neural NetworksAndrea Bonfanti, Giuseppe Bruno, Cristina CiprianiNeurIPS 2024 · 被引用 41 次
