Decoupled Greedy Learning of CNNs
Eugene Belilovsky, Michael Eickenberg, Edouard Oyallon
摘要
A commonly cited inefficiency of neural network training by back-propagation is the update locking problem: each layer must wait for the signal to propagate through the full network before updating. Several alternatives that can alleviate this issue have been proposed. In this context, we consider a simpler, but more effective, substitute that uses minimal feedback, which we call Decoupled Greedy Learning (DGL). It is based on a greedy relaxation of the joint training objective, recently shown to be effective in the context of Convolutional Neural Networks (CNNs) on large-scale image classification. We consider an optimization of this objective that permits us to decouple the layer training, allowing for layers or modules in networks to be trained with a potentially linear parallelization in layers. With the use of a replay buffer we show this approach can be extended to asynchronous settings, where modules can operate with poossibly large communication delays. We show theoretically and empirically that this approach converges. Then, we empirically find that it can lead to better generalization than sequential greedy optimization. We demonstrate the effectiveness of DGL against alternative approaches on the CIFAR-10 dataset and on the large-scale ImageNet dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Revisiting Locally Supervised Learning: an Alternative to End-to-end TrainingYulin Wang, Zanlin Ni, Shiji Song, Le Yang 等ICLR 2021 · 被引用 99 次
- Online Learned Continual Compression with Adaptive Quantization ModulesLucas Caccia, Eugene Belilovsky, Massimo Caccia, Joelle PineauICML 2020 · 被引用 95 次
- Error-driven Input Modulation: Solving the Credit Assignment Problem without a Backward PassGiorgia Dellaferrera, Gabriel KreimanICML 2022 · 被引用 80 次
- The staircase property: How hierarchical structure can guide deep learningEmmanuel Abbe, Enric Boix-Adserà, Matthew S. Brennan, Guy Bresler 等NeurIPS 2021 · 被引用 74 次
- ProgFed: Effective, Communication, and Computation Efficient Federated Learning by Progressive TrainingHui-Po Wang, Sebastian U. Stich, Yang He, Mario FritzICML 2022 · 被引用 70 次
相关 Paper
- Accumulated Decoupled Learning with Gradient Staleness Mitigation for Convolutional Neural NetworksHuiping Zhuang, Zhenyu Weng, Fulin Luo, Kar-Ann Toj 等ICML 2021 · 被引用 6 次
- Backpropagation-Free Deep Learning with Recursive Local Representation AlignmentAlexander G. Ororbia II, Ankur Mali, Daniel Kifer, C. Lee GilesAAAI 2023 · 被引用 19 次
- SEDONA: Search for Decoupled Neural Networks toward Greedy Block-wise LearningMyeongjang Pyeon, Jihwan Moon, Taeyoung Hahn, Gunhee KimICLR 2021 · 被引用 14 次
- Scaling Supervised Local Learning with Augmented Auxiliary NetworksChenxiang Ma, Jibin Wu, Chenyang Si, Kay Chen TanICLR 2024 · 被引用 9 次
- Towards Interpretable Deep Local Learning with Successive Gradient ReconciliationYibo Yang, Xiaojie Li, Motasem Alfarra, Hasan Abed Al Kader Hammoud 等ICML 2024 · 被引用 8 次
