Subway: minimizing data transfer during out-of-GPU-memory graph processing
Amir Hossein Nodehi Sabet, Zhijia Zhao, Rajiv Gupta
摘要
In many graph-based applications, the graphs tend to grow, imposing a great challenge for GPU-based graph processing. When the graph size exceeds the device memory capacity (i.e., GPU memory oversubscription), the performance of graph processing often degrades dramatically, due to the sheer amount of data transfer between CPU and GPU.
To reduce the volume of data transfer, existing approaches track the activeness of graph partitions and only load the ones that need to be processed. In fact, the recent advances of unified memory implements this optimization implicitly by loading memory pages on demand. However, either way, the benefits are limited by the coarse-granularity activeness tracking -each loaded partition or memory page may still carry a large ratio of inactive edges.
In this work, we present, to the best of our knowledge, the first solution that only loads active edges of the graph to the GPU memory. To achieve this, we design a fast subgraph generation algorithm with a simple yet efficient subgraph representation and a GPU-accelerated implementation. They allow the subgraph generation to be applied in almost every iteration of the vertex-centric graph processing. Furthermore, we bring asynchrony to the subgraph processing, delaying the synchronization between a subgraph in the GPU memory and the rest of the graph in the CPU memory. This can safely reduce the needs of generating and loading subgraphs for many common graph algorithms. Our prototyped system, Subway (subgraph processing with asynchrony) yields over 4X speedup on average comparing with existing out-of-GPUmemory solutions and the unified memory-based approach, based on an evaluation with six common graph algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Large Graph Convolutional Network Training with GPU-Oriented Data Communication ArchitectureSeungwon Min, Kun Wu, Sitao Huang, Mert Hidayetoglu 等VLDB 2021 · 被引用 85 次
- Accelerating graph sampling for graph machine learning using GPUsAbhinav Jangda, Sandeep Polisetty, Arjun Guha, Marco SerafiniEuroSys 2021 · 被引用 79 次
- EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal In GPUsSeungwon Min, Vikram Sharma Mailthody, Zaid Qureshi, Jinjun Xiong 等VLDB 2021 · 被引用 66 次
- CommonGraph: Graph Analytics on Evolving DataMahbod Afarin, Chao Gao, Shafiur Rahman, Nael B. Abu-Ghazaleh 等ASPLOS 2023 · 被引用 32 次
- HongTu: Scalable Full-Graph GNN Training on Multiple GPUsQiange Wang, Yao Chen, Weng-Fai Wong, Bingsheng HeSIGMOD 2024 · 被引用 24 次
相关 Paper
- CGgraph: An Ultra-fast Graph Processing System on Modern Commodity CPU-GPU Co-processorPengjie Cui, Haotian Liu, Bo Tang, Ye YuanVLDB 2024 · 被引用 18 次
- HyTGraph: GPU-Accelerated Graph Processing with Hybrid Transfer ManagementQiange Wang, Xin Ai, Yanfeng Zhang, Jing Chen 等ICDE 2023 · 被引用 14 次
- GPU-Accelerated Subgraph Enumeration on Partitioned GraphsWentian Guo, Yuchen Li, Mo Sha, Bingsheng He 等SIGMOD 2020 · 被引用 71 次
- GPU-Accelerated Batch-Dynamic Subgraph MatchingLinshan Qiu, Lu Chen, Hailiang Jie, Xiangyu Ke 等ICDE 2024 · 被引用 7 次
- LightTraffic: On Optimizing CPU-GPU Data Traffic for Efficient Large-scale Random WalksYipeng Xing, Yongkun Li, Zhiqiang Wang, Yinlong Xu 等ICDE 2023 · 被引用 3 次
