Lune

ICDE2025Top-tier venue

CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access

Ziyu Song, Jie Zhang, Jie Sun, Mo Sun, Zihan Yang, Zheng Zhang, Xuzheng Chen, Fei Wu, Huajin Tang, Zeke Wang

2025Year
4Citations
2Top-tier citations

Abstract

With the wide adoption of GPU and the explosion in data volumes, existing accelerator-centric systems require massive storage access. They adopt high-performance storage devices like NVMe SSDs to scale up single-node systems costeffectively and leverage the CPU to manage these SSDs. However, they suffer from performance bottlenecks because of the high CPU OS kernel overhead and the CPU memory intermediated data transfer. To address this issue, GPU-initiated and GPUmanaged SSD management is proposed to allow the GPU to fully manipulate SSDs: 1) direct data transfer from SSD to GPU memory (data plane) and 2) GPU-managed SSD control (control plane). This can potentially enable these GPU systems to fully leverage the SSD bandwidth. However, we still identify two severe issues. First, the GPU-management SSD control leads to low GPU Streaming Multiprocessor utilization. Second, it leads to the serial execution of SSD accesses with GPU computation, which slows down the overall computing task. To this end, we propose CAM, the first asynchronous GPU-initialized, CPU-managed SSD management for batching storage access. It 1) offloads the SSD control plane from GPU to CPU, thus maximizing GPU streaming multiprocessor utilization, and 2) adopts asynchronous user-friendly APIs that allow programmers to easily overlap GPU computation and SSD I/O operations while keeping a synchronous programming experience. As such, CAM enables us to achieve the best of two worlds: high performance and high programmability. The experimental results show that CAM can perform GNN model training, mergesort, and GEMM up to 1.84×, 1.5×, and 1.84× faster, compared to the existing state-ofthe-art GPU systems, while keeping high programmability.

With the advancement of GPUs, many cutting-edge applications, such as neural network models [1], [21] and GPUbased database systems [6], [50], are turning into GPU-centric systems, which can benefit from GPU's massive parallel computing power. In particular, the NVIDIA A100 GPU delivers 312 TeraFLOPS (TFLOPS) of computing capability, while the AMD Threadripper 3995WX CPU has fewer than 3 TFLOPS. Together with the increasing requirement of computing power, the problem size of an application also increases faster. For GNN, the graph can contain billions of vertices and tens of billions of edges [35], [59], [63], which needs several terabytes of storage space. For DLRM, the memory capacity of embedding tables has increased dramatically from tens of GBs to TBs throughout the industry [33], [61], [62]. Therefore, many researches [43], [46], [48], [58] leverage SSDs to break the GPU memory and server memory boundaries so as to enable out-of-core computation on massive data volume for a broad range of applications. Storing data in SSDs not only 2309

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 863a554e-03e3-4010-8fc4-eb31ce68ab31

Cited by top-tier papers2

Ask how each one uses it

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines