EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal In GPUs
Seungwon Min, Vikram Sharma Mailthody, Zaid Qureshi, Jinjun Xiong, Eiman Ebrahimi, Wen-Mei Hwu
摘要
Modern analytics and recommendation systems are increasingly based on graph data that capture the relations between entities being analyzed. Practical graphs come in huge sizes, offer massive parallelism, and are stored in sparse-matrix formats such as compressed sparse row (CSR). To exploit the massive parallelism, developers are increasingly interested in using GPUs for graph traversal. However, due to their sizes, graphs often do not fit into the GPU memory. Prior works have either used input data pre-processing/partitioning or unified virtual memory (UVM) to migrate chunks of data from the host memory to the GPU memory. However, the large, multi-dimensional, and sparse nature of graph data presents a major challenge to these schemes and results in significant amplification of data movement and reduced effective data throughput. In this work, we propose EMOGI, an alternative approach to traverse graphs that do not fit in GPU memory using direct cache-line-sized access to data stored in host memory. This paper addresses the open question of whether a sufficiently large number of overlapping cache-line-sized accesses can be sustained to 1) tolerate the long latency to host memory, 2) fully utilize the available bandwidth, and 3) achieve favorable execution performance. We analyze the data access patterns of several graph traversal applications in GPU over PCIe using an FPGA to understand the cause of poor external bandwidth utilization. By carefully coalescing and aligning external memory requests, we show that we can minimize the number of PCIe transactions and nearly fully utilize the PCIe bandwidth with direct cache-line accesses to the host memory. EMOGI achieves 2.60X speedup on average compared to the optimized UVM implementations in various graph traversal applications. We also show that EMOGI scales better than a UVM-based solution when the system uses higher bandwidth interconnects such as PCIe 4.0.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Large Graph Convolutional Network Training with GPU-Oriented Data Communication ArchitectureSeungwon Min, Kun Wu, Sitao Huang, Mert Hidayetoglu 等VLDB 2021 · 被引用 85 次
- In-depth analyses of unified virtual memory system for GPU accelerated computingTyler N. Allen, Rong GeSC 2021 · 被引用 73 次
- RecShard: statistical feature-based memory optimization for industry-scale neural recommendationGeet Sethi, Bilge Acun, Niket Agarwal, Christos Kozyrakis 等ASPLOS 2022 · 被引用 65 次
- GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System ArchitectureZaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado, Seungwon Min 等ASPLOS 2023 · 被引用 48 次
- Strata: Hierarchical Context Caching for Long Context Language Model ServingZhiqiang Xie, Ziyi Xu, Mark Zhao, Yuwei An 等OSDI 2026 · 被引用 40 次
它引用的顶会 Paper3
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi 等ASPLOS 2020 · 被引用 89 次
- Subway: minimizing data transfer during out-of-GPU-memory graph processingAmir Hossein Nodehi Sabet, Zhijia Zhao, Rajiv GuptaEuroSys 2020 · 被引用 84 次
- Traversing Large Graphs on GPUs with Unified MemoryPrasun Gera, Hyojong Kim, Piyush Sao, Hyesoon Kim 等VLDB 2020 · 被引用 58 次
相关 Paper
- Random Walks on Huge Graphs at Cache EfficiencyKe Yang, Xiaosong Ma, Saravanan Thirumuruganathan, Kang Chen 等SOSP 2021 · 被引用 26 次
- Self-adaptive Graph Traversal on GPUsMo Sha, Yuchen Li, Kian-Lee TanSIGMOD 2021 · 被引用 12 次
- Efficient GPU-Accelerated Subgraph MatchingXibo Sun, Qiong LuoSIGMOD 2023 · 被引用 29 次
- HyTGraph: GPU-Accelerated Graph Processing with Hybrid Transfer ManagementQiange Wang, Xin Ai, Yanfeng Zhang, Jing Chen 等ICDE 2023 · 被引用 14 次
- GFlux: A Fast GPU-Based Out-of-Memory Multi-Hop Query Processing Framework for Trillion-Edge GraphsSeyeon Oh, Heeyong Yoon, Donghyoung Han, Min-Soo KimICDE 2025 · 被引用 1 次
