FlashNeuron: SSD-Enabled Large-Batch Training of Very Deep Neural Networks
Jonghyun Bae, Jongsung Lee, Yunho Jin, Sam Son, Shine Kim, Hakbeom Jang, Tae Jun Ham, Jae W. Lee
摘要
© 2021 by The USENIX Association.Deep neural networks (DNNs) are widely used in various AI application domains such as computer vision, natural language processing, autonomous driving, and bioinformatics. As DNNs continue to get wider and deeper to improve accuracy, the limited DRAM capacity of a training platform like GPU often becomes the limiting factor on the size of DNNs and batch size—called memory capacity wall. Since increasing the batch size is a popular technique to improve hardware utilization, this can yield a suboptimal training throughput. Recent proposals address this problem by offloading some of the intermediate data (e.g., feature maps) to the host memory. However, they fail to provide robust performance as the training process on a GPU contends with applications running on a CPU for memory bandwidth and capacity. Thus, we propose FlashNeuron, the first DNN training system using an NVMe SSD as a backing store. To fully utilize the limited SSD write bandwidth, FlashNeuron introduces an offloading scheduler, which selectively offloads a set of intermediate data to the SSD in a compressed format without increasing DNN evaluation time. FlashNeuron causes minimal interference to CPU processes as the GPU and the SSD directly communicate for data transfers. Our evaluation of FlashNeuron with four state-of-the-art DNNs shows that FlashNeuron can increase the batch size by a factor of 12.4× to 14.0× over the maximum allowable batch size on NVIDIA Tesla V100 GPU with 16GB DRAM. By employing a larger batch size, FlashNeuron also improves the training throughput by up to 37.8% (with an average of 30.3%) over the baseline using GPU memory only, while minimally disturbing applications running on CPU.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory CachingYeonhong Park, Sunhong Min, Jae W. LeeVLDB 2022 · 被引用 57 次
- NVMeVirt: A Versatile Software-defined Virtual NVMe DeviceSang-Hoon Kim, Jaehoon Shim, Euidong Lee, Seong-Yeob Jeong 等FAST 2023 · 被引用 52 次
- Mobius: Fine Tuning Large-Scale Models on Commodity GPU ServersYangyang Feng, Minhui Xie, Zijie Tian, Shuo Wang 等ASPLOS 2023 · 被引用 29 次
- Fast State Restoration in LLM Serving with HCacheShiwei Gao, Youmin Chen, Jiwu ShuEuroSys 2025 · 被引用 22 次
- G10: Enabling An Efficient Unified GPU Memory and Storage Architecture with Smart Tensor MigrationsHaoyang Zhang, Yirui Eric Zhou, Yuqi Xue, Yiqi Liu 等MICRO 2023 · 被引用 21 次
它引用的顶会 Paper4
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- Analyzing and Mitigating Data Stalls in DNN TrainingJayashree Mohan, Amar Phanishayee, Ashish Raniwala, Vijay ChidambaramVLDB 2021 · 被引用 142 次
- Quiver: An Informed Storage Cache for Deep LearningAbhishek Vijaya Kumar, Muthian SivathanuFAST 2020 · 被引用 91 次
- Echo: Compiler-based GPU Memory Footprint Reduction for LSTM RNN TrainingBojian Zheng, Nandita Vijaykumar, Gennady PekhimenkoISCA 2020 · 被引用 34 次
相关 Paper
- OptimStore: In-Storage Optimization of Large Scale DNNs with On-Die ProcessingJunkyum Kim, Myeonggu Kang, Yunki Han, Yang-gon Kim 等HPCA 2023 · 被引用 25 次
- STRONGHOLD: Fast and Affordable Billion-Scale Deep Learning Model TrainingXiaoyang Sun, Wei Wang, Shenghao Qiu, Renyu Yang 等SC 2022 · 被引用 17 次
- Behemoth: A Flash-centric Training Accelerator for Extreme-scale DNNsShine Kim, Yunho Jin, Gina Sohn, Jonghyun Bae 等FAST 2021 · 被引用 44 次
- SSDTrain: An Activation Offloading Framework to SSDs for Faster Large Language Model TrainingKun Wu, Jeongmin Brian Park, Xiaofan Zhang, Mert Hidayetoglu 等DAC 2025 · 被引用 3 次
- FlashGNN: An In-SSD Accelerator for GNN TrainingFuping Niu, Jianhui Yue, Jiangqiu Shen, Xiaofei Liao 等HPCA 2024 · 被引用 13 次
