A theory of weight distribution-constrained learning
Weishun Zhong, Ben Sorscher, Daniel Lee, Haim Sompolinsky
Abstract
A central question in computational neuroscience is how structure determines function in neural networks. Recent large-scale connectomic studies have started to provide a wealth of structural information such as the distribution of excitatory/inhibitory cell and synapse types as well as the distribution of synaptic weights in the brains of different species. The emerging high-quality large structural datasets raise the question of what general functional principles can be gleaned from them. Motivated by this question, we developed a statistical mechanical theory of learning in neural networks that incorporates structural information as constraints. We derived an analytical solution for the memory capacity of the perceptron, a basic feedforward model of supervised learning, with constraint on the distribution of its weights. Interestingly, the theory predicts that the reduction in capacity due to the constrained weight-distribution is related to the Wasserstein distance between the cumulative distribution function of the constrained weights and that of the standard normal distribution. To test the theoretical predictions, we use optimal transport theory and information geometry to develop an SGDbased algorithm to find weights that simultaneously learn the input-output task and satisfy the distribution constraint. We show that training in our algorithm can be interpreted as geodesic flows in the Wasserstein space of probability distributions. We further developed a statistical mechanical theory for teacher-student perceptron rule learning and ask for the best way for the student to incorporate prior knowledge of the rule (i.e., the teacher). Our theory shows that it is beneficial for the learner to adopt different prior weight distributions during learning, and shows that distribution-constrained learning outperforms unconstrained and sign-constrained learning. Our theory and algorithm provide novel strategies for incorporating prior knowledge about weights into learning, and reveal a powerful connection between structure and function in neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d388736-4b88-41a6-a185-eb9a29ba79cdCited by top-tier papers1
Ask how each one uses itBuilds on2
- Learning curves of generic features maps for realistic datasets with a teacher-student modelBruno Loureiro, Cédric Gerbelot, Hugo Cui, Sebastian Goldt et al.NeurIPS 2021 · 170 citations
- Continual Learning in the Teacher-Student Setup: Impact of Task SimilaritySebastian Lee, Sebastian Goldt, Andrew M. SaxeICML 2021 · 98 citations
Related papers
- The impact of allocation strategies in subset learning on the expressive power of neural networksOfir Schlisselberg, Ran DarshanICLR 2025
- Biologically plausible heavy-tailed connectivity enhances generalizations on cognitive tasks in recurrent neural networksZhe Jiao, Xiaodong He, Shanglin ZhouICML 2026
- Synaptic Weight Distributions Depend on the Geometry of PlasticityRoman Pogodin, Jonathan Cornford, Arna Ghosh, Gauthier Gidel et al.ICLR 2024 · 8 citations
- Learning Useful Representations of Recurrent Neural Network Weight MatricesVincent Herrmann, Francesco Faccio, Jürgen SchmidhuberICML 2024 · 12 citations
- Spike-based causal inference for weight alignmentJordan Guerguiev, Konrad P. Körding, Blake A. RichardsICLR 2020 · 26 citations
