AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
Zhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones, Peipei Zhou
摘要
GPUs are critical for compute-intensive applications, yet emerging workloads such as recommender systems, graph analytics, and data analytics often exceed GPU memory capacity. Existing solutions allow GPUs to use CPU DRAM or SSDs as external memory, and the GPU-centric approach enables GPU threads to directly issue NVMe requests, further avoiding CPU intervention. However, current GPU-centric approaches adopt synchronous I/O, forcing threads to stall during long communication delays.
We propose AGILE, a lightweight asynchronous GPU-centric I/O library that eliminates deadlock risks and integrates a flexible HBM-based software cache. AGILE overlaps computation and I/O, improving performance by up to 1.88× across workloads with diverse computation-to-communication ratios. Compared to BaM on DLRM, AGILE achieves up to 1.75× speedup through efficient design and overlapping; on graph applications, AGILE reduces software cache overhead by up to 3.12× and NVMe I/O overhead by up to 2.85×; AGILE also lowers per-thread register usage by up to 1.32×.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Overcoming the Memory Wall with CXL-Enabled SSDsShao-Peng Yang, Minjae Kim, Sanghyun Nam, Juhyung Park 等USENIX ATC 2023 · 被引用 75 次
- In-depth analyses of unified virtual memory system for GPU accelerated computingTyler N. Allen, Rong GeSC 2021 · 被引用 73 次
- FpgaNIC: An FPGA-based Versatile 100Gb SmartNIC for GPUsZeke Wang, Hongjing Huang, Jie Zhang, Fei Wu 等USENIX ATC 2022 · 被引用 58 次
- Centauri: Enabling Efficient Scheduling for Communication-Computation Overlap in Large Model Training via Communication PartitioningChang Chen, Xiuhong Li, Qianchao Zhu, Jiangfei Duan 等ASPLOS 2024 · 被引用 52 次
- GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System ArchitectureZaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado, Seungwon Min 等ASPLOS 2023 · 被引用 48 次
相关 Paper
- Asynchrony and GPUs: Bridging this Dichotomy for I/O with AGIOJihoon Han, Anand Sivasubramaniam, Chia-Hao Chang, Vikram Sharma Mailthody 等ASPLOS 2026 · 被引用 1 次
- CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage AccessZiyu Song, Jie Zhang, Jie Sun, Mo Sun 等ICDE 2025 · 被引用 4 次
- GeminiFS: A Companion File System for GPUsShi Qiu, Weinan Liu, Yifan Hu, Jianqin Yan 等FAST 2025 · 被引用 17 次
- Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic OperationsYicong Zhang, Mingyu Wang, Wangguang Wang, Yangzhan Mai 等MICRO 2024 · 被引用 4 次
- MDM: The GPU Memory Divergence ModelLu Wang, Magnus Jahre, Almutaz Adileh, Lieven EeckhoutMICRO 2020 · 被引用 27 次
