Towards Interpretable Deep Local Learning with Successive Gradient Reconciliation
Yibo Yang, Xiaojie Li, Motasem Alfarra, Hasan Abed Al Kader Hammoud, Adel Bibi, Philip Torr, Bernard Ghanem
Abstract
Relieving the reliance of neural network training on a global back-propagation (BP) has emerged as a notable research topic due to the biological implausibility and huge memory consumption caused by BP. Among the existing solutions, local learning optimizes gradient-isolated modules of a neural network with local errors and has been proved to be effective even on large-scale datasets. However, the reconciliation among local errors has never been investigated. In this paper, we first theoretically study non-greedy layer-wise training and show that the convergence cannot be assured when the local gradient in a module w.r.t. its input is not reconciled with the local gradient in the previous module w.r.t. its output. Inspired by the theoretical result, we further propose a local training strategy that successively regularizes the gradient reconciliation between neighboring modules without breaking gradient isolation or introducing any learnable parameters. Our method can be integrated into both local-BP and BP-free settings. In experiments, we achieve significant performance improvements compared to previous methods. Particularly, our method for CNN and Transformer architectures on ImageNet is able to attain a competitive performance with global BP, saving more than 40% memory consumption.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-SpaceBingjie Zhang, Yibo Yang, Renzhe, Dandan Guo et al.ICLR 2026 · 12 citations
- CorDA: Context-Oriented Decomposition Adaptation of Large Language Models for Task-Aware Parameter-Efficient Fine-tuningYibo Yang, Xiaojie Li, Zhongzhu Zhou, Shuaiwen Song et al.NeurIPS 2024
Builds on17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li et al.NeurIPS 2021 · 303 citations
- Inducing Neural Collapse in Imbalanced Learning: Do We Really Need a Learnable Classifier at the End of Deep Neural Network?Yibo Yang, Shixiang Chen, Xiangtai Li, Liang Xie et al.NeurIPS 2022 · 144 citations
Related papers
- Module-wise Training of Neural Networks via the Minimizing Movement SchemeSkander Karkar, Ibrahim Ayed, Emmanuel de Bézenac, Patrick GallinariNeurIPS 2023 · 6 citations
- Scaling Supervised Local Learning with Augmented Auxiliary NetworksChenxiang Ma, Jibin Wu, Chenyang Si, Kay Chen TanICLR 2024 · 9 citations
- Decoupled Greedy Learning of CNNsEugene Belilovsky, Michael Eickenberg, Edouard OyallonICML 2020 · 134 citations
- Backpropagation-Free Deep Learning with Recursive Local Representation AlignmentAlexander G. Ororbia II, Ankur Mali, Daniel Kifer, C. Lee GilesAAAI 2023 · 19 citations
- Forward Learning of Graph Neural NetworksNamyong Park, Xing Wang, Antoine Simoulin, Shuai Yang et al.ICLR 2024 · 1 citation
