GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System Architecture
Zaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado, Seungwon Min, Amna Masood, Jeongmin Brian Park, Jinjun Xiong, Chris J. Newburn, Dmitri Vainbrand, I-Hsin Chung, Michael Garland, William J. Dally, Wen-Mei W. Hwu
摘要
Graphics Processing Units (GPUs) have traditionally relied on the host CPU to initiate access to the data storage. This approach is well-suited for GPU applications with known data access patterns that enable partitioning of their dataset to be processed in a pipelined fashion in the GPU. However, emerging applications such as graph and data analytics, recommender systems, or graph neural networks, require fine-grained, data-dependent access to storage. CPU initiation of storage access is unsuitable for these applications due to high CPU-GPU synchronization overheads, I/O traffic amplification, and long CPU processing latencies. GPU-initiated storage removes these overheads from the storage control path and, thus, can potentially support these applications at much higher speed. However, there is a lack of systems architecture and software stack that enable efficient GPU-initiated storage access. This work presents a novel system architecture, BaM, that fills this gap. BaM features a fine-grained software cache to coalesce data storage requests while minimizing I/O traffic amplification. This software cache communicates with the storage system via high-throughput queues that enable the massive number of concurrent threads in modern GPUs to make I/O requests at a high rate to fully utilize the storage devices and the system interconnect. Experimental results show that BaM delivers 1.0x and 1.49x end-to-end speed up for BFS and CC graph analytics benchmarks while reducing hardware costs by up to 21.7x over accessing the graph data from the host memory. Furthermore, BaM speeds up data-analytics workloads by 5.3x over CPU-initiated storage access on the same hardware.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Strata: Hierarchical Context Caching for Long Context Language Model ServingZhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An 等OSDI 2026 · 被引用 40 次
- Accelerating Sampling and Aggregation Operations in GNN Frameworks with GPU Initiated Direct Storage AccessesJeongmin Brian Park, Vikram Sharma Mailthody, Zaid Qureshi, Wen-Mei HwuVLDB 2024 · 被引用 37 次
- M5: Mastering Page Migration and Memory Management for CXL-based Tiered Memory SystemsYan Sun, Jongyul Kim, Zeduo Yu, Jiyuan Zhang 等ASPLOS 2025 · 被引用 27 次
- Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real SystemHongsun Jang, Jaeyong Song, Jaewon Jung, Jaeyoung Park 等HPCA 2024 · 被引用 26 次
- Beluga: A CXL-Based Memory Architecture for Scalable and Efficient LLM KVCache ManagementXinjun Yang, Qingda Hu, Junru Li, Feifei Li 等SIGMOD 2026 · 被引用 24 次
它引用的顶会 Paper8
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith 等SC 2021 · 被引用 254 次
- LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismJongyul Kim, Insu Jang, Waleed Reda, Jaeseong Im 等SOSP 2021 · 被引用 83 次
- EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal In GPUsSeungwon Min, Vikram Sharma Mailthody, Zaid Qureshi, Jinjun Xiong 等VLDB 2021 · 被引用 66 次
- Lynx: A SmartNIC-driven Accelerator-centric Architecture for Network ServersMaroun Tork, Lina Maudlej, Mark SilbersteinASPLOS 2020 · 被引用 64 次
- CrossFS: A Cross-layered Direct-Access File SystemYujie Ren, Changwoo Min, Sudarsun KannanOSDI 2020 · 被引用 35 次
相关 Paper
- AGILE: Lightweight and Efficient Asynchronous GPU-SSD IntegrationZhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones 等SC 2025 · 被引用 3 次
- CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage AccessZiyu Song, Jie Zhang, Jie Sun, Mo Sun 等ICDE 2025 · 被引用 4 次
- GeminiFS: A Companion File System for GPUsShi Qiu, Weinan Liu, Yifan Hu, Jianqin Yan 等FAST 2025 · 被引用 17 次
- CoPilotIO: CPU as a Co-Pilot for GPU I/O to Free GPU ComputeGuanyi Chen, Qi Chen, Shu Yin, Jian ZhangOSDI 2026
- GLIST: Towards In-Storage Graph LearningCangyuan Li, Ying Wang, Cheng Liu, Shengwen Liang 等USENIX ATC 2021 · 被引用 53 次
