On the Convergence of Two-Layer Kolmogorov-Arnold Networks with First-Layer Training
Seyed Mohammad Eshtehardian, Mohammad Hossein Yassaee, Babak HosseinKhalaj
摘要
Kolmogorov-Arnold Networks (KANs) have emerged as a promising alternative to traditional neural networks, offering enhanced interpretability based on the Kolmogorov-Arnold representation theorem. While their empirical success is growing, a theoretical understanding of their training dynamics remains nascent. This paper investigates the optimization of a two-layer KAN in the overparameterized regime, focusing on a simplified yet insightful setting where only the first-layer coefficients are trained via gradient descent.
Our main result establishes that, provided the network is sufficiently wide, this training method is guaranteed to converge to a global minimum and achieve zero training error. Furthermore, we derive a novel, fine-grained convergence rate that explicitly connects the optimization speed to the structure of the data labels through the eigenspectrum of the KAN Tangent Kernel (KAN-TK). Our analysis reveals a key advantage of this architecture: guaranteed convergence is achieved with a hidden layer width of , a significant polynomial improvement over the requirement for classic two-layer neural networks using ReLU activation functions and analyzed within the same Tangent Kernel framework. We validate our theoretical findings with numerical experiments that corroborate our predictions on convergence speed and the impact of label structure.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- U-KAN Makes Strong Backbone for Medical Image Segmentation and GenerationChenxin Li, Xinyu Liu, Wuyang Li, Cheng Wang 等AAAI 2025 · 被引用 452 次
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 被引用 193 次
- Initialization Schemes for Kolmogorov–Arnold Networks: An Empirical StudySpyros Rigas, Dhruv Verma, Georgios Alexandridis, Yixuan WangICLR 2026 · 被引用 13 次
- KAA: Kolmogorov-Arnold Attention for Enhancing Attentive Graph Neural NetworksTaoran Fang, Tianhong Gao, Chunping Wang, Yihao Shang 等ICLR 2025 · 被引用 1 次
- Generalization Bounds and Model Complexity for Kolmogorov-Arnold NetworksXianyang Zhang, Huijuan ZhouICLR 2025
相关 Paper
- Optimization, Generalization and Differential Privacy Bounds for Gradient Descent on Kolmogorov–Arnold NetworksPuyu Wang, Junyu Zhou, Philipp Liznerski, Marius KloftICML 2026 · 被引用 2 次
- On feature learning in neural networks with global convergence guaranteesZhengdao Chen, Eric Vanden-Eijnden, Joan BrunaICLR 2022 · 被引用 15 次
- Optimal Rates for Averaged Stochastic Gradient Descent under Neural Tangent Kernel RegimeAtsushi Nitanda, Taiji SuzukiICLR 2021 · 被引用 49 次
- Mean-field Analysis on Two-layer Neural Networks from a Kernel PerspectiveShokichi Takakura, Taiji SuzukiICML 2024 · 被引用 12 次
- PowerMLP: An Efficient Version of KANRuichen Qiu, Yibo Miao, Shiwen Wang, Yifan Zhu 等AAAI 2025 · 被引用 13 次
