SC2025Top-tier venue
AGILE: Lightweight and Efficient Asynchronous GPU-SSD Integration
Zhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones, Peipei Zhou
Abstract
GPUs are critical for compute-intensive applications, yet emerging workloads such as recommender systems, graph analytics, and data analytics often exceed GPU memory capacity. Existing solutions allow GPUs to use CPU DRAM or SSDs as external memory, and the GPU-centric approach enables GPU threads to directly issue NVMe requests, further avoiding CPU intervention. However, current GPU-centric approaches adopt synchronous I/O, forcing threads to stall during long communication delays.
We propose AGILE, a lightweight asynchronous GPU-centric I/O library that eliminates deadlock risks and integrates a flexible HBM-based software cache. AGILE overlaps computation and I/O, improving performance by up to 1.88× across workloads with diverse computation-to-communication ratios. Compared to BaM on DLRM, AGILE achieves up to 1.75× speedup through efficient design and overlapping; on graph applications, AGILE reduces software cache overhead by up to 3.12× and NVMe I/O overhead by up to 2.85×; AGILE also lowers per-thread register usage by up to 1.32×.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98b62a5d-441a-4329-843e-952e79525a52Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Overcoming the Memory Wall with CXL-Enabled SSDsShao-Peng Yang, Minjae Kim, Sanghyun Nam, Juhyung Park et al.USENIX ATC 2023 · 75 citations
- In-depth analyses of unified virtual memory system for GPU accelerated computingTyler N. Allen, Rong GeSC 2021 · 73 citations
- FpgaNIC: An FPGA-based Versatile 100Gb SmartNIC for GPUsZeke Wang, Hongjing Huang, Jie Zhang, Fei Wu et al.USENIX ATC 2022 · 58 citations
- Centauri: Enabling Efficient Scheduling for Communication-Computation Overlap in Large Model Training via Communication PartitioningChang Chen, Xiuhong Li, Qianchao Zhu, Jiangfei Duan et al.ASPLOS 2024 · 52 citations
- GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System ArchitectureZaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado, Seungwon Min et al.ASPLOS 2023 · 48 citations
Related papers
- Asynchrony and GPUs: Bridging this Dichotomy for I/O with AGIOJihoon Han, Anand Sivasubramaniam, Chia-Hao Chang, Vikram Sharma Mailthody et al.ASPLOS 2026 · 1 citation
- CAM: Asynchronous GPU-Initiated, CPU-Managed SSD Management for Batching Storage AccessZiyu Song, Jie Zhang, Jie Sun, Mo Sun et al.ICDE 2025 · 4 citations
- GeminiFS: A Companion File System for GPUsShi Qiu, Weinan Liu, Yifan Hu, Jianqin Yan et al.FAST 2025 · 17 citations
- Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic OperationsYicong Zhang, Mingyu Wang, Wangguang Wang, Yangzhan Mai et al.MICRO 2024 · 4 citations
- MDM: The GPU Memory Divergence ModelLu Wang, Magnus Jahre, Almutaz Adileh, Lieven EeckhoutMICRO 2020 · 27 citations
