Brain network science modelling of sparse neural networks enables Transformers and LLMs to perform as fully connected
Yingtao Zhang, Diego Cerretti, Jialin Zhao, Wenjing Wu, Ziheng Liao, Umberto Michieli, Carlo Vittorio Cannistraci
Abstract
This study aims to enlarge our current knowledge on the application of braininspired network science principles for training artificial neural networks (ANNs) with sparse connectivity. Dynamic sparse training (DST) emulates the synaptic turnover of real brain networks, reducing the computational demands of training and inference in ANNs. However, existing DST methods face difficulties in maintaining peak performance at high connectivity sparsity levels. The Cannistraci-Hebb training (CHT) is a brain-inspired method that is used in DST for growing synaptic connectivity in sparse neural networks. CHT leverages a gradient-free, topology-driven link regrowth mechanism, which has been shown to achieve ultrasparse (1% connectivity or lower) advantage across various tasks compared to fully connected networks. Yet, CHT suffers two main drawbacks: (i) its time complexity is O(N • d 3 )-N node network size, d node degree -hence it can be efficiently applied only to ultra-sparse networks. (ii) it rigidly selects top link prediction scores, which is inappropriate for the early training epochs, when the network topology presents many unreliable connections. Here, we design the first brain-inspired network model -termed bipartite receptive field (BRF) -to initialize the connectivity of sparse artificial neural networks. Then, we propose a matrix multiplication GPU-friendly approximation of the CH link predictor, which reduces the computational complexity to O(N 3 ), enabling a fast implementation of link prediction in large-scale models. Moreover, we introduce the Cannistraci-Hebb training soft rule (CHTs), which adopts a flexible strategy for sampling connections in both link removal and regrowth, balancing the exploration and exploitation of network topology. Additionally, we propose a sigmoid-based gradual density decay strategy, leading to an advanced framework referred to as CHTss. Empirical results show that BRF offers performance advantages over previous network science models. Using 1% of connections, CHTs outperforms fully connected networks in MLP architectures on visual classification tasks, compressing some networks to less than 30% of the nodes. Using 5% of the connections, CHTss outperforms fully connected networks in two Transformer-based machine translation tasks. Finally, with only 30% of the connections, both CHTs and CHTss achieve superior performance over other dynamic sparse training methods, and perform on par with-or even surpass-their fully connected counterparts in language modeling across various sparsity levels within the LLaMA model family.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d60219c5-43f6-442d-95ec-642a48fbfc7dCited by top-tier papers3
- Cannistraci-Hebb Training on Ultra-Sparse Spiking Neural NetworksYuan Hua, Jilin Zhang, Yingtao Zhang, Leyi You et al.ICLR 2026 · 2 citations
- Sign-In to the Lottery: Reparameterizing Sparse TrainingAdvait Gadhikar, Tom Jacobs, Chao Zhou, Rebekka BurkholzNeurIPS 2025
- Alignment-Enhanced Integration of Connectivity and Spectral Sparsity in Dynamic Sparse Training of LLMWenjing Wu, Yingtao Zhang, Jialin Zhao, Carlo Vittorio CannistraciICLR 2026
Builds on21
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank ProjectionJiawei Zhao, Zhenyu Zhang, Beidi Chen, Zhangyang Wang et al.ICML 2024 · 433 citations
- DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMsHaokun Lin, Haobo Xu, Yichen Wu, Jingzhi Cui et al.NeurIPS 2024 · 206 citations
Related papers
- Epitopological learning and Cannistraci-Hebb network shape intelligence brain-inspired theory for ultra-sparse advantage in deep learningYingtao Zhang, Jialin Zhao, Wenjing Wu, Alessandro Muscoloni et al.ICLR 2024 · 10 citations
- Adaptive Cannistraci-Hebb Network Automata Modelling of Complex Networks for Path-based Link PredictionJialin Zhao, Alessandro Muscoloni, Umberto Michieli, Yingtao Zhang et al.NeurIPS 2025 · 1 citation
- Fantastic Weights and How to Find Them: Where to Prune in Dynamic Sparse TrainingAleksandra Nowak, Bram Grooten, Decebal Constantin Mocanu, Jacek TaborNeurIPS 2023 · 23 citations
- ESL-SNNs: An Evolutionary Structure Learning Strategy for Spiking Neural NetworksJiangrong Shen, Qi Xu, Jian K. Liu, Yueming Wang et al.AAAI 2023 · 64 citations
- Neurogenesis Dynamics-inspired Spiking Neural Network Training AccelerationShaoyi Huang, Haowen Fang, Kaleel Mahmood, Bowen Lei et al.DAC 2023 · 6 citations
