NeuroFlux: Memory-Efficient CNN Training Using Adaptive Local Learning
Dhananjay Saikumar, Blesson Varghese
摘要
Efficient on-device Convolutional Neural Network (CNN) training in resource-constrained mobile and edge environments is an open challenge. Backpropagation is the standard approach adopted, but it is GPU memory intensive due to its strong inter-layer dependencies that demand intermediate activations across the entire CNN model to be retained in GPU memory. This necessitates smaller batch sizes to make training possible within the available GPU memory budget, but in turn, results in substantially high and impractical training time. We introduce NeuroFlux, a novel CNN training system tailored for memory-constrained scenarios. We develop two novel opportunities: firstly, adaptive auxiliary networks that employ a variable number of filters to reduce GPU memory usage, and secondly, block-specific adaptive batch sizes, which not only cater to the GPU memory constraints but also accelerate the training process. NeuroFlux segments a CNN into blocks based on GPU memory usage and further attaches an auxiliary network to each layer in these blocks. This disrupts the typical layer dependencies under a new training paradigm - 'adaptive local learning'. Moreover, NeuroFlux adeptly caches intermediate activations, eliminating redundant forward passes over previously trained blocks, further accelerating the training process. The results are twofold when compared to Backpropagation: on various hardware platforms, NeuroFlux demonstrates training speed-ups of 2.3× to 6.1× under stringent GPU memory budgets, and NeuroFlux generates streamlined models that have 10.9× to 29.4× fewer parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev 等ICLR 2020 · 被引用 229 次
- Improved Techniques for Training Adaptive Deep NetworksHao Li, Hong Zhang, Xiaojuan Qi, Ruigang Yang 等ICCV 2019 · 被引用 152 次
- Decoupled Greedy Learning of CNNsEugene Belilovsky, Michael Eickenberg, Edouard OyallonICML 2020 · 被引用 134 次
- REFL: Resource-Efficient Federated LearningAhmed M. Abdelmoniem, Atal Narayan Sahu, Marco Canini, Suhaib A. FahmyEuroSys 2023 · 被引用 86 次
相关 Paper
- Efficient On-Device Training via Gradient FilteringYuedong Yang, Guihong Li, Radu MarculescuCVPR 2023
- Enabling On-Tiny-Device Model Personalization via Gradient Condensing and Alternant Partial UpdateZhenge Jia, Yiyang Shi, Zeyu Bao, Zirui Wang 等DAC 2025
- MLAAN: Scaling Supervised Local Learning with Multilaminar Leap Augmented Auxiliary NetworkYuming Zhang, Shouxin Zhang, Peizhe Wang, Feiyu Zhu 等AAAI 2025 · 被引用 4 次
- FreshGNN: Reducing Memory Access via Stable Historical Embeddings for Graph Neural Network TrainingKezhao Huang, Haitian Jiang, Minjie Wang, Guangxuan Xiao 等VLDB 2024 · 被引用 13 次
- Group Knowledge Transfer: Federated Learning of Large CNNs at the EdgeChaoyang He, Murali Annavaram, Salman AvestimehrNeurIPS 2020 · 被引用 605 次
