Initialization Schemes for Kolmogorov–Arnold Networks: An Empirical Study
Spyros Rigas, Dhruv Verma, Georgios Alexandridis, Yixuan Wang
摘要
Kolmogorov–Arnold Networks (KANs) are a recently introduced neural architecture that replace fixed nonlinearities with trainable activation functions, offering enhanced flexibility and interpretability. While KANs have been applied successfully across scientific and machine learning tasks, their initialization strategies remain largely unexplored. In this work, we study initialization schemes for spline-based KANs, proposing two theory-driven approaches inspired by LeCun and Glorot, as well as an empirical power-law family with tunable exponents. Our evaluation combines large-scale grid searches on function fitting and forward PDE benchmarks, an analysis of training dynamics through the lens of the Neural Tangent Kernel, and evaluations on a subset of the Feynman dataset. Our findings indicate that the Glorot-inspired initialization significantly outperforms the baseline in parameter-rich models, while power-law initialization achieves the strongest performance overall, both across tasks and for architectures of varying size. This work underscores initialization as a key factor in KAN performance and introduces practical strategies to improve it.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
- Improving Transformer Optimization Through Better InitializationXiao Shi Huang, Felipe Pérez, Jimmy Ba, Maksims VolkovsICML 2020 · 被引用 181 次
- KANO: Kolmogorov-Arnold Neural OperatorJin Lee, Ziming Liu, Xinling Yu, Yixuan Wang 等ICLR 2026 · 被引用 6 次
- Generalization Bounds and Model Complexity for Kolmogorov-Arnold NetworksXianyang Zhang, Huijuan ZhouICLR 2025
- Deep Learning Alternatives Of The Kolmogorov Superposition TheoremLeonardo Ferreira Guilhoto, Paris PerdikarisICLR 2025
相关 Paper
- KAN: Kolmogorov-Arnold NetworksZiming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle 等ICLR 2025
- PowerMLP: An Efficient Version of KANRuichen Qiu, Yibo Miao, Shiwen Wang, Yifan Zhu 等AAAI 2025 · 被引用 13 次
- Incorporating Arbitrary Matrix Group Equivariance into KANsLexiang Hu, Yisen Wang, Zhouchen LinICML 2025
- U-KAN Makes Strong Backbone for Medical Image Segmentation and GenerationChenxin Li, Xinyu Liu, Wuyang Li, Cheng Wang 等AAAI 2025 · 被引用 452 次
- On the expressiveness and spectral bias of KANsYixuan Wang, Jonathan W. Siegel, Ziming Liu, Thomas Y. HouICLR 2025
