CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage Access
Ziyu Song, Jie Zhang, Jie Sun, Mo Sun, Zihan Yang, Zheng Zhang, Xuzheng Chen, Fei Wu, Huajin Tang, Zeke Wang
Abstract
With the wide adoption of GPU and the explosion in data volumes, existing accelerator-centric systems require massive storage access. They adopt high-performance storage devices like NVMe SSDs to scale up single-node systems costeffectively and leverage the CPU to manage these SSDs. However, they suffer from performance bottlenecks because of the high CPU OS kernel overhead and the CPU memory intermediated data transfer. To address this issue, GPU-initiated and GPUmanaged SSD management is proposed to allow the GPU to fully manipulate SSDs: 1) direct data transfer from SSD to GPU memory (data plane) and 2) GPU-managed SSD control (control plane). This can potentially enable these GPU systems to fully leverage the SSD bandwidth. However, we still identify two severe issues. First, the GPU-management SSD control leads to low GPU Streaming Multiprocessor utilization. Second, it leads to the serial execution of SSD accesses with GPU computation, which slows down the overall computing task. To this end, we propose CAM, the first asynchronous GPU-initialized, CPU-managed SSD management for batching storage access. It 1) offloads the SSD control plane from GPU to CPU, thus maximizing GPU streaming multiprocessor utilization, and 2) adopts asynchronous user-friendly APIs that allow programmers to easily overlap GPU computation and SSD I/O operations while keeping a synchronous programming experience. As such, CAM enables us to achieve the best of two worlds: high performance and high programmability. The experimental results show that CAM can perform GNN model training, mergesort, and GEMM up to 1.84×, 1.5×, and 1.84× faster, compared to the existing state-ofthe-art GPU systems, while keeping high programmability.
With the advancement of GPUs, many cutting-edge applications, such as neural network models [1], [21] and GPUbased database systems [6], [50], are turning into GPU-centric systems, which can benefit from GPU's massive parallel computing power. In particular, the NVIDIA A100 GPU delivers 312 TeraFLOPS (TFLOPS) of computing capability, while the AMD Threadripper 3995WX CPU has fewer than 3 TFLOPS. Together with the increasing requirement of computing power, the problem size of an application also increases faster. For GNN, the graph can contain billions of vertices and tens of billions of edges [35], [59], [63], which needs several terabytes of storage space. For DLRM, the memory capacity of embedding tables has increased dramatically from tens of GBs to TBs throughout the industry [33], [61], [62]. Therefore, many researches [43], [46], [48], [58] leverage SSDs to break the GPU memory and server memory boundaries so as to enable out-of-core computation on massive data volume for a broad range of applications. Storing data in SSDs not only 2309
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 863a554e-03e3-4010-8fc4-eb31ce68ab31Cited by top-tier papers2
- Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN TrainingZuocheng Shi, Jie Sun, Ziyu Song, Mo Sun et al.SC 2025 · 3 citations
- CoPilotIO: CPU as a Co-Pilot for GPU I/O to Free GPU ComputeGuanyi Chen, Qi Chen, Shu Yin, Jian ZhangOSDI 2026
Builds on21
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith et al.SC 2021 · 254 citations
- What Modern NVMe Storage Can Do, And How To Exploit It: High-Performance I/O for High-Performance Storage EnginesGabriel Haas, Viktor LeisVLDB 2023 · 83 citations
- Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory CachingYeonhong Park, Sunhong Min, Jae W. LeeVLDB 2022 · 57 citations
- ScaleStore: A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMATobias Ziegler, Carsten Binnig, Viktor LeisSIGMOD 2022 · 51 citations
Related papers
- Managing Scalable Direct Storage Accesses for GPUs with GoFSShaobo Li, Yirui Eric Zhou, Yuqi Xue, Yuan Xu et al.SOSP 2025
- GeminiFS: A Companion File System for GPUsShi Qiu, Weinan Liu, Yifan Hu, Jianqin Yan et al.FAST 2025 · 17 citations
- Hyperion: Co-Optimizing SSD Access and GPU Computation for Cost-Efficient GNN TrainingJie Sun, Mo Sun, Zheng Zhang, Zuocheng Shi et al.ICDE 2025 · 3 citations
- GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System ArchitectureZaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado, Seungwon Min et al.ASPLOS 2023 · 48 citations
- AGILE: Lightweight and Efficient Asynchronous GPU-SSD IntegrationZhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones et al.SC 2025 · 3 citations
