Depth-Progressive Monotonic Learning without Global Backpropagation
Chenhao Ye, Rongguang Ye, Yuchao Zhang, Ming Tang
Abstract
Backpropagation (BP), while foundational to deep learning, imposes two critical scalability bottlenecks: update locking, where network modules remain idle until the entire backward pass completes, and high memory consumption due to storing activations for gradient computation. To address these limitations, we introduce Synergistic Information Distillation (SID), a novel training framework that reframes deep learning as a cascade of local cooperative refinement problems. In SID, a deep network is structured as a pipeline of modules, each imposed with a local objective to refine a probabilistic "belief" about the ground-truth target. This objective balances fidelity to the target with consistency to the belief from its preceding module. By decoupling the backward dependencies between modules, SID enables parallel training and hence eliminates update locking and drastically reduces memory requirements. Meanwhile, this design preserves the standard feed-forward inference pass, making SID a versatile drop-in replacement for BP. We provide a theoretical foundation, proving that SID guarantees monotonic performance improvement with network depth. Empirically, SID consistently matches or surpasses the classification accuracy of BP, exhibiting superior scalability and pronounced robustness to label noise. The code is publicly available at: https://github.com/ychAlbert/sid_bp .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbe538d3-256f-41d3-aa4d-11055be8c8abBuilds on8
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- The HSIC Bottleneck: Deep Learning without Back-PropagationKurt Wan-Duo Ma, J. P. Lewis, W. Bastiaan KleijnAAAI 2020 · 180 citations
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang et al.NeurIPS 2024 · 50 citations
- Towards Scaling Difference Target Propagation by Learning Backprop TargetsMaxence Ernoult, Fabrice Normandin, Abhinav Moudgil, Sean Spinney et al.ICML 2022 · 49 citations
Related papers
- Scaling Supervised Local Learning with Augmented Auxiliary NetworksChenxiang Ma, Jibin Wu, Chenyang Si, Kay Chen TanICLR 2024 · 9 citations
- Sideways: Depth-Parallel Training of Video ModelsMateusz Malinowski, Grzegorz Swirszcz, João Carreira, Viorica PatrauceanCVPR 2020
- Decoupled Greedy Learning of CNNsEugene Belilovsky, Michael Eickenberg, Edouard OyallonICML 2020 · 134 citations
- Accelerated training through iterative gradient propagation along the residual pathErwan Fagnou, Paul Caillon, Blaise Delattre, Alexandre AllauzenICLR 2025
- Enabling Efficient Large-Scale Deep Learning Training with Cache Coherent Disaggregated Memory SystemsZixuan Wang, Joonseop Sim, Euicheol Lim, Jishen ZhaoHPCA 2022 · 9 citations
