Kolmogorov-Arnold Networks Still Catastrophically Forget but Differently from MLP
Anton Lee, Heitor Murilo Gomes, Yaqian Zhang, W. Bastiaan Kleijn
Abstract
Catastrophic forgetting is when a neural network loses previously learnt information after learning a new task sequentially. Avoiding catastrophic forgetting could reduce the resources necessary to update neural networks. Recently, Kolmogorov-Arnold Networks (KAN) gained the community's attention as preliminary experiments suggest KAN avoid catastrophic forgetting. KAN replace neural network edges with learnable B-splines and sum incoming edges in nodes. Proponents of KAN argue they avoid forgetting, are more accurate, are interpretable, and use fewer parameters. Our work investigates the claims that KAN avoid catastrophic forgetting, finding that they fail to do so on more complex datasets containing features that overlap between tasks. We give a simple explanation as to why and how KAN catastrophically forget. Motivated by evidence suggesting KAN are superior for symbolic regression, we augment KAN in the same ways as multilayer perceptron (MLP) to perform continual learning tasks, making special accommodations to support KAN. Our experiments found that unmodified KAN often forget more than MLP, but KAN can be better than MLP when combined with continual learning strategies. We aim to highlight some of the current shortcomings and strengths associated with KAN for continual learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Unifying Locality of KANs and Feature Drift Compensation Projection for Data-Free Replay Based Continual Face Forgery DetectionTianshuo Zhang, Siran Peng, Li Gao, Haoyuan Zhang et al.AAAI 2026 · 1 citation
- Catastrophic Forgetting in Kolmogorov-Arnold NetworksMohammad Marufur Rahman, Guanchu Wang, Kaixiong Zhou, Minghan Chen et al.AAAI 2026 · 1 citation
Builds on3
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati et al.NeurIPS 2020 · 1,494 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
- Wide Neural Networks Forget Less CatastrophicallySeyed-Iman Mirzadeh, Arslan Chaudhry, Dong Yin, Huiyi Hu et al.ICML 2022 · 84 citations
Related papers
- KAC: Kolmogorov-Arnold Classifier for Continual LearningYusong Hu, Zichen Liang, Fei Yang, Qibin Hou et al.CVPR 2025
- Neuro-Symbolic Continual Learning: Knowledge, Reasoning Shortcuts and Concept RehearsalEmanuele Marconato, Gianpaolo Bontempo, Elisa Ficarra, Simone Calderara et al.ICML 2023 · 34 citations
- KAN: Kolmogorov-Arnold NetworksZiming Liu, Yixuan Wang, Sachin Vaidya, Fabian Ruehle et al.ICLR 2025
- PowerMLP: An Efficient Version of KANRuichen Qiu, Yibo Miao, Shiwen Wang, Yifan Zhu et al.AAAI 2025 · 13 citations
- Learning curves for continual learning in neural networks: Self-knowledge transfer and forgettingRyo Karakida, Shotaro AkahoICLR 2022 · 16 citations
