Bandwidth-Effective DRAM Cache for GPU s with Storage-Class Memory
Jeongmin Hong, Sungjun Cho, Geonwoo Park, Wonhyuk Yang, Young-Ho Gong, Gwangsun Kim
摘要
We propose overcoming the memory capacity limitation of GPUs with high-capacity Storage-Class Memory (SCM) and DRAM cache. By significantly increasing the memory capacity with SCM, the GPU can capture a larger fraction of the memory footprint than HBM for workloads that mandate memory oversubscription, resulting in substantial speedups. However, the DRAM cache needs to be carefully designed to address the latency and bandwidth limitations of the SCM while minimizing cost overhead and considering GPU's characteristics. Because the massive number of GPU threads can easily thrash the DRAM cache and degrade performance, we first propose an SCM-aware DRAM cache bypass policy for GPUs that considers the multi-dimensional characteristics of memory accesses by G PU s with SCM to bypass DRAM for data with low performance utility. In addition, to reduce DRAM cache probe traffic and increase effective DRAM BW with minimal cost overhead, we propose a Configurable Tag Cache (CTC) that repurposes part of the L2 cache to cache DRAM cacheline tags. The L2 capacity used for the CTC can be adjusted by users for adaptability. Furthermore, to minimize DRAM cache probe traffic from CTC misses, our Aggregated Metadata-In-Last-column (AMIL) DRAM cache organization co-locates all DRAM cacheline tags in a single column within a row. The AMIL also retains the full ECC protection, unlike prior DRAM cache implementation with Tag-And-Data (TAD) organization. Additionally, we propose SCM throttling to curtail power consumption and exploiting SCM's SLC/MLC modes to adapt to workload's memory footprint. While our techniques can be used for different DRAM and SCM devices, we focus on a Heterogeneous Memory Stack (HMS) organization that stacks SCM dies on top of DRAM dies for high performance. Compared to HBM, the HMS improves performance by up to 12.5× (2.9× overall) and reduces energy by up to 89.3% (48.1 % overall). Compared to prior works, we reduce DRAM cache probe and SCM write traffic by 91–93 % and 57–75 %, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FlexNeRFer: A Multi-Dataflow, Adaptive Sparsity-Aware Accelerator for On-Device NeRF RenderingSeock-Hwan Noh, Banseok Shin, Jeik Choi, Seungpyo Lee 等ISCA 2025 · 被引用 3 次
- AGILE: Lightweight and Efficient Asynchronous GPU-SSD IntegrationZhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones 等SC 2025 · 被引用 3 次
- Efficient Caching with A Tag-enhanced DRAMMaryam Babaie, Ayaz Akram, Wendy Elsasser, Brent Haukness 等HPCA 2025 · 被引用 2 次
- Xerxes: Extensive Exploration of Scalable Hardware Systems with CXL-Based Simulation FrameworkYuda An, Shushu Yi, Bo Mao, Qiao Li 等FAST 2026 · 被引用 2 次
它引用的顶会 Paper21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 被引用 366 次
- AccelWattch: A Power Modeling Framework for Modern GPUsVijay Kandiah, Scott Peverelle, Mahmoud Khairy, Junrui Pan 等MICRO 2021 · 被引用 134 次
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi 等ASPLOS 2020 · 被引用 89 次
相关 Paper
- Native DRAM Cache: Re-architecting DRAM as a Large-Scale Cache for Data CentersYesin Ryu, Yoojin Kim, Giyong Jung, Jung Ho Ahn 等ISCA 2024 · 被引用 5 次
- GMT: GPU Orchestrated Memory Tiering for the Big Data EraChia-Hao Chang, Jihoon Han, Anand Sivasubramaniam, Vikram Sharma Mailthody 等ASPLOS 2024 · 被引用 11 次
- Observability-Aided Gpu Memory OversubscriptionPratheek B, Khushit Shah, Arkaprava BasuISCA 2026
- NOMAD: Enabling Non-blocking OS-managed DRAM Cache via Tag-Data DecouplingYoungin Kim, Hyeonjin Kim, William J. SongHPCA 2023 · 被引用 10 次
- ZnG: Architecting GPU Multi-Processors with New Flash for Scalable Data AnalysisJie Zhang, Myoungsoo JungISCA 2020 · 被引用 14 次
