Behemoth: A Flash-centric Training Accelerator for Extreme-scale DNNs
Shine Kim, Yunho Jin, Gina Sohn, Jonghyun Bae, Tae Jun Ham, Jae W. Lee
摘要
© 2021 by The USENIX Association.The explosive expansion of Deep Neural Networks (DNN) model size expedites the need for larger memory capacity. This movement is particularly true for models in natural language processing (NLP), a dominant application of AI along with computer vision. For example, a recent extreme-scale language model GPT-3 from OpenAI has over 175 billion parameters. Furthermore, such a model mostly consists of FC layers with huge dimensions, and thus has a relatively high arithmetic intensity. In that sense, an extreme-scale language model does not suit well to the conventional HBM DRAM-based memory system that lacks capacity and offers extremely high bandwidth. For this reason, we propose to pair the neural network training accelerator with the flash-based memory system instead of the HBM DRAM-based memory system. To design the effective flash-based memory system, we optimize the existing SSD design to improve the SSD bandwidth as well as endurance. Finally, we evaluate our proposed platform, and show that Behemoth achieves 3.65× cost saving over TPU v3 and 2.05× training throughput improvement over the accelerator attached to a commercial SSD.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Overcoming the Memory Wall with CXL-Enabled SSDsShao-Peng Yang, Minjae Kim, Sanghyun Nam, Juhyung Park 等USENIX ATC 2023 · 被引用 75 次
- Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory CachingYeonhong Park, Sunhong Min, Jae W. LeeVLDB 2022 · 被引用 57 次
- λ-IO: A Unified IO Stack for Computational StorageZhe Yang, Youyou Lu, Xiaojian Liao, Youmin Chen 等FAST 2023 · 被引用 54 次
- Flash-Cosmos: In-Flash Bulk Bitwise Operations Using Inherent Computation Capability of NAND Flash MemoryJisung Park, Roknoddin Azizi, Geraldo F. Oliveira, Mohammad Sadrosadati 等MICRO 2022 · 被引用 53 次
- Cambricon-LLM: A Chiplet-Based Hybrid Architecture for On-Device Inference of 70B LLMZhongkai Yu, Shengwen Liang, Tianyun Ma, Yunke Cai 等MICRO 2024 · 被引用 29 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- Layerweaver: Maximizing Resource Utilization of Neural Processing Units via Layer-Wise SchedulingYoung H. Oh, Seonghak Kim, Yunho Jin, Sam Son 等HPCA 2021 · 被引用 46 次
相关 Paper
- FlashNeuron: SSD-Enabled Large-Batch Training of Very Deep Neural NetworksJonghyun Bae, Jongsung Lee, Yunho Jin, Sam Son 等FAST 2021 · 被引用 64 次
- IANUS: Integrated Accelerator based on NPU-PIM Unified Memory SystemMinseok Seo, Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon 等ASPLOS 2024 · 被引用 57 次
- H3T: Efficient Integration of Memory Optimization and Parallelism for Large-scale Transformer TrainingYuzhong Wang, Xu Han, Weilin Zhao, Guoyang Zeng 等NeurIPS 2023 · 被引用 2 次
- TrainBox: An Extreme-Scale Neural Network Training Server Architecture by Systematically Balancing OperationsPyeongsu Park, Heetaek Jeong, Jangwoo KimMICRO 2020 · 被引用 11 次
- TransPIM: A Memory-based Acceleration via Software-Hardware Co-Design for TransformerMinxuan Zhou, Weihong Xu, Jaeyoung Kang, Tajana RosingHPCA 2022 · 被引用 142 次
