Optimization, Generalization and Differential Privacy Bounds for Gradient Descent on Kolmogorov–Arnold Networks
Puyu Wang, Junyu Zhou, Philipp Liznerski, Marius Kloft
摘要
Kolmogorov--Arnold Networks (KANs) have recently emerged as a structured alternative to standard MLPs, yet a principled theory for their training dynamics, generalization, and privacy properties remains limited. In this paper, we analyze gradient descent (GD) for training two-layer KANs and derive general bounds that characterize their training dynamics, generalization, and utility under differential privacy (DP). As a concrete instantiation, we specialize our analysis to logistic loss under an NTK-separable assumption, where we show that polylogarithmic network width suffices for GD to achieve an optimization rate of order and a generalization rate of order , with denoting the number of GD iterations and the sample size. In the private setting, we characterize the noise required for -DP and obtain a utility bound of order (with the input dimension), matching the classical lower bound for general convex Lipschitz problems. Our results imply that polylogarithmic width is not only sufficient but also necessary under differential privacy, revealing a qualitative gap between non-private (sufficiency only) and private (necessity also emerges) training regimes. Experiments further illustrate how these theoretical insights can guide practical choices, including network width selection and early stopping.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 被引用 193 次
- Fine-Grained Analysis of Stability and Generalization for Stochastic Gradient DescentYunwen Lei, Yiming YingICML 2020 · 被引用 165 次
- Differentially Private Stochastic Optimization: New Results in Convex and Non-Convex SettingsRaef Bassily, Cristóbal Guzmán, Michael MenartNeurIPS 2021 · 被引用 68 次
- Optimal Rates for Averaged Stochastic Gradient Descent under Neural Tangent Kernel RegimeAtsushi Nitanda, Taiji SuzukiICLR 2021 · 被引用 49 次
相关 Paper
- On the Convergence of Two-Layer Kolmogorov-Arnold Networks with First-Layer TrainingSeyed Mohammad Eshtehardian, Mohammad Hossein Yassaee, Babak HosseinKhalajICLR 2026
- Generalization Bounds and Model Complexity for Kolmogorov-Arnold NetworksXianyang Zhang, Huijuan ZhouICLR 2025
- On the expressiveness and spectral bias of KANsYixuan Wang, Jonathan W. Siegel, Ziming Liu, Thomas Y. HouICLR 2025
- Beyond NTK with Vanilla Gradient Descent: A Mean-Field Analysis of Neural Networks with Polynomial Width, Samples, and TimeArvind V. Mahankali, Haochen Zhang, Kefan Dong, Margalit Glasgow 等NeurIPS 2023 · 被引用 20 次
- Sharper Guarantees for Learning Neural Network Classifiers with Gradient MethodsHossein Taheri, Christos Thrampoulidis, Arya MazumdarICLR 2025
