MLAAN: Scaling Supervised Local Learning with Multilaminar Leap Augmented Auxiliary Network
Yuming Zhang, Shouxin Zhang, Peizhe Wang, Feiyu Zhu, Dongzhi Guan, Junhao Su, Jiabin Liu, Changpeng Cai
Abstract
Deep neural networks (DNNs) typically employ an end-to-end (E2E) training paradigm which presents several challenges, including high GPU memory consumption, inefficiency, and difficulties in model parallelization during training. Recent research has sought to address these issues, with one promising approach being local learning. This method involves partitioning the backbone network into gradient-isolated modules and manually designing auxiliary networks to train these local modules. Existing methods often neglect the interaction of information between local modules, leading to myopic issues and a performance gap compared to E2E training. To address these limitations, we propose the Multilaminar Leap Augmented Auxiliary Network (MLAAN). Specifically, MLAAN comprises Multilaminar Local Modules (MLM) and Leap Augmented Modules (LAM). MLM captures both local and global features through independent and cascaded auxiliary networks, alleviating performance issues caused by insufficient global features. However, overly simplistic auxiliary networks can impede MLM's ability to capture global information. To address this, we further design LAM, an enhanced auxiliary network that uses the Exponential Moving Average (EMA) method to facilitate information exchange between local modules, thereby mitigating the shortsightedness resulting from inadequate interaction. The synergy between MLM and LAM has demonstrated excellent performance. Our experiments on the CIFAR-10, STL-10, SVHN, and ImageNet datasets show that MLAAN can be seamlessly integrated into existing local learning frameworks, significantly enhancing their performance and even surpassing end-to-end (E2E) training methods, while also reducing GPU memory consumption.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext debc7800-8f53-421a-8773-6a3f7e1bf7a1Builds on5
- Decoupled Greedy Learning of CNNsEugene Belilovsky, Michael Eickenberg, Edouard OyallonICML 2020 · 134 citations
- Revisiting Locally Supervised Learning: an Alternative to End-to-end TrainingYulin Wang, Zanlin Ni, Shiji Song, Le Yang et al.ICLR 2021 · 99 citations
- Local plasticity rules can learn deep representations using self-supervised contrastive predictionsBernd Illing, Jean Ventura, Guillaume Bellec, Wulfram GerstnerNeurIPS 2021 · 99 citations
- Error-driven Input Modulation: Solving the Credit Assignment Problem without a Backward PassGiorgia Dellaferrera, Gabriel KreimanICML 2022 · 80 citations
- Focus on Local: Detecting Lane Marker From Bottom Up via Key PointZhan Qu, Huan Jin, Yang Zhou, Zhen Yang et al.CVPR 2021
Related papers
- Scaling Supervised Local Learning with Augmented Auxiliary NetworksChenxiang Ma, Jibin Wu, Chenyang Si, Kay Chen TanICLR 2024 · 9 citations
- NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local LearningDhananjay Saikumar, Blesson VargheseEuroSys 2024 · 2 citations
- Towards Interpretable Deep Local Learning with Successive Gradient ReconciliationYibo Yang, Xiaojie Li, Motasem Alfarra, Hasan Abed Al Kader Hammoud et al.ICML 2024 · 8 citations
- Efficient GPU Memory Management for Nonlinear DNNsDonglin Yang, Dazhao ChengHPDC 2020 · 18 citations
- Capuchin: Tensor-based GPU Memory Management for Deep LearningXuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin et al.ASPLOS 2020 · 143 citations
