GMT: GPU Orchestrated Memory Tiering for the Big Data Era
Chia-Hao Chang, Jihoon Han, Anand Sivasubramaniam, Vikram Sharma Mailthody, Zaid Qureshi, Wen-Mei Hwu
Abstract
As the demand for processing larger datasets increases, GPUs need to reach deeper into their (memory) hierarchy to directly access capacities that only storage systems (SSDs) can hold. However, the state-of-the-art mechanisms to reach storage either employ software stacks running on the host CPUs as intermediaries (e.g. Dragon, HMM), which has been noted to perform poorly and not able to meet the throughput needs of GPU cores, or directly access SSDs through NVMe queues (BaM) which does not benefit from lower latencies that may be possible by having the host memory as an intermediate tier. This paper presents the design and implementation of GPU Memory Tiering (GMT) by implementing a GPU-orchestrated 3-tier hierarchy comprising GPU memory, host memory and SSDs, where the GPU orchestrates most of the transfers that are bandwidth/latency sensitive. Additionally, it is important to not blindly transfer pages from the GPU memory to host memory upon an eviction, and GMT employs a reuse-prediction based practical insertion policy to perform discretionary page placement/bypass. An implementation and evaluation on an actual platform demonstrates that GMT performs 50% better than the state-of-the-art 2-tier strategy (BaM) and over 350% better than the state-of-the-art 3-tier strategy that is orchestrated by host CPUs (HMM), over a number of GPU applications with diverse memory access characteristics.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 9131820e-9ff9-4ab2-9b40-1bc1b17904fdCited by top-tier papers4
- GeminiFS: A Companion File System for GPUsShi Qiu, Weinan Liu, Yifan Hu, Jianqin Yan et al.FAST 2025 · 17 citations
- AGILE: Lightweight and Efficient Asynchronous GPU-SSD IntegrationZhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones et al.SC 2025 · 3 citations
- CoPilotIO: CPU as a Co-Pilot for GPU I/O to Free GPU ComputeGuanyi Chen, Qi Chen, Shu Yin, Jian ZhangOSDI 2026
- A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMsHongsun Jang, Jaeyong Song, Changmin Shin, Si Ung Noh et al.ASPLOS 2026
Related papers
- Automating Distributed Tiered Storage Management in Cluster ComputingHerodotos Herodotou, Elena KakoulliVLDB 2020 · 30 citations
- Bandwidth-Effective DRAM Cache for GPU s with Storage-Class MemoryJeongmin Hong, Sungjun Cho, Geonwoo Park, Wonhyuk Yang et al.HPCA 2024 · 21 citations
- ZnG: Architecting GPU Multi-Processors with New Flash for Scalable Data AnalysisJie Zhang, Myoungsoo JungISCA 2020 · 14 citations
- MegaMmap: Blurring the Boundary Between Memory and Storage for Data-Intensive WorkloadsLuke Logan, Anthony Kougkas, Xian-He SunSC 2024 · 3 citations
- Managing Scalable Direct Storage Accesses for GPUs with GoFSShaobo Li, Yirui Eric Zhou, Yuqi Xue, Yuan Xu et al.SOSP 2025
