Buddy Compression: Enabling Larger Memory for Deep Learning and HPC Workloads on GPUs
Esha Choukse, Michael B. Sullivan, Mike O'Connor, Mattan Erez, Jeff Pool, David W. Nellans, Stephen W. Keckler
摘要
GPUs accelerate high-throughput applications, which require orders-of-magnitude higher memory bandwidth than traditional CPU-only systems. However, the capacity of such high-bandwidth memory tends to be relatively small. Buddy Compression is an architecture that makes novel use of compression to utilize a larger buddy-memory from the host or disaggregated memory, effectively increasing the memory capacity of the GPU. Buddy Compression splits each compressed 128B memory-entry between the high-bandwidth GPU memory and a slower-but-larger buddy memory such that compressible memory-entries are accessed completely from GPU memory, while incompressible entries source some of their data from off-GPU memory. With Buddy Compression, compressibility changes never result in expensive page movement or re-allocation. Buddy Compression achieves on average 1.9× effective GPU memory expansion for representative HPC applications and 1.5× for deep learning training, performing within 2% of an unrealistic system with no memory limit. This makes Buddy Compression attractive for performance-conscious developers that require additional GPU memory capacity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
- Toward Sustainable HPC: Carbon Footprint Estimation and Environmental Implications of HPC SystemsBaolin Li, Rohan Basu Roy, Daniel Wang, Siddharth Samsi 等SC 2023 · 被引用 72 次
- Zico: Efficient GPU Memory Sharing for Concurrent DNN TrainingGangmuk Lim, Jeongseob Ahn, Wencong Xiao, Youngjin Kwon 等USENIX ATC 2021 · 被引用 65 次
- Tensor Casting: Co-Designing Algorithm-Architecture for Personalized Recommendation TrainingYoungeun Kwon, Yunjae Lee, Minsoo RhuHPCA 2021 · 被引用 40 次
- COMET: A Novel Memory-Efficient Deep Learning Training Framework by Using Error-Bounded Lossy CompressionSian Jin, Chengming Zhang, Xintong Jiang, Yunhe Feng 等VLDB 2022 · 被引用 39 次
相关 Paper
- Tile-based Lightweight Integer Compression in GPUAnil Shanbhag, Bobbi W. Yogatama, Xiangyao Yu, Samuel MaddenSIGMOD 2022 · 被引用 45 次
- FZ-GPU: A Fast and High-Ratio Lossy Compressor for Scientific Computing Applications on GPUsBoyuan Zhang, Jiannan Tian, Sheng Di, Xiaodong Yu 等HPDC 2023 · 被引用 27 次
- cuSZp2: A GPU Lossy Compressor with Extreme Throughput and Optimized Compression RatioYafan Huang, Sheng Di, Guanpeng Li, Franck CappelloSC 2024 · 被引用 29 次
- cuSZp: An Ultra-fast GPU Error-bounded Lossy Compression Framework with Optimized End-to-End PerformanceYafan Huang, Sheng Di, Xiaodong Yu, Guanpeng Li 等SC 2023 · 被引用 52 次
- Optimizing Random Access to Hierarchically-Compressed Data on GPUFeng Zhang, Yihua Hu, Haipeng Ding, Zhiming Yao 等SC 2022 · 被引用 5 次
