How gradient estimator variance and bias impact learning in neural networks
Arna Ghosh, Yuhan Helena Liu, Guillaume Lajoie, Konrad P. Körding, Blake Aaron Richards
摘要
There is growing interest in understanding how real brains may approximate gradients and how gradients can be used to train neuromorphic chips. However, neither real brains nor neuromorphic chips can perfectly follow the loss gradient, so parameter updates would necessarily use gradient estimators that have some variance and/or bias. Therefore, there is a need to understand better how variance and bias in gradient estimators impact learning dependent on network and task properties. Here, we show that variance and bias can impair learning on the training data, but some degree of variance and bias in a gradient estimator can be beneficial for generalization. We find that the ideal amount of variance and bias in a gradient estimator are dependent on several properties of the network and task: the size and activity sparsity of the network, the norm of the gradient, and the curvature of the loss landscape. As such, whether considering biologically-plausible learning algorithms or algorithms for training neuromorphic chips, researchers can analyze these properties to determine whether their approximation to gradient descent will be effective for learning given their network and task properties.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- How connectivity structure shapes rich and lazy learning in neural circuitsYuhan Helena Liu, Aristide Baratin, Jonathan Cornford, Stefan Mihalas 等ICLR 2024 · 被引用 26 次
- Dynamic Momentum Recalibration in Online Gradient LearningZhipeng Yao, Rui Yu, Guisong Chang, Ying Li 等CVPR 2026 · 被引用 1 次
- Can Biologically Plausible Temporal Credit Assignment Rules Match BPTT for Neural Similarity? E-prop as an ExampleYuhan Helena Liu, Guangyu Robert Yang, Christopher J. CuevaICML 2025
相关 Paper
- Beyond accuracy: generalization properties of bio-plausible temporal credit assignment rulesYuhan Helena Liu, Arna Ghosh, Blake A. Richards, Eric Shea-Brown 等NeurIPS 2022 · 被引用 10 次
- Penalising the biases in norm regularisation enforces sparsityEtienne Boursier, Nicolas FlammarionNeurIPS 2023 · 被引用 21 次
- Improving equilibrium propagation without weight symmetry through Jacobian homeostasisAxel Laborieux, Friedemann ZenkeICLR 2024 · 被引用 11 次
- Path Sample-Analytic Gradient Estimators for Stochastic Binary NetworksAlexander Shekhovtsov, Viktor Yanush, Boris FlachNeurIPS 2020 · 被引用 14 次
- Adaptive Smoothing Gradient Learning for Spiking Neural NetworksZiming Wang, Runhao Jiang, Shuang Lian, Rui Yan 等ICML 2023 · 被引用 69 次
