Two-level Graph Caching for Expediting Distributed GNN Training
Zhe Zhang, Ziyue Luo, Chuan Wu
Abstract
Graph Neural Networks (GNNs) are increasingly popular due to excellent performance on learning graphstructured data in various domains. With fast expanding graph sizes and feature dimensions, distributed GNN training has been adopted, with multiple concurrent workers learning on different portions of a large graph. It has been observed that a main bottleneck in distributed GNN training lies in graph feature fetching across servers, which dominates the training time of each training iteration at each worker. This paper studies efficient feature caching on each worker to minimize feature fetching overhead, in order to expedite distributed GNN training. Current distributed GNN training systems largely adopt static caching of fixed neighbor nodes. We propose a novel two-level dynamic cache design exploiting both GPU memory and host memory at each worker, and design efficient two-level dynamic caching algorithms based on online optimization and a lookahead batching mechanism. Our dynamic caching algorithms consider node requesting probabilities and heterogeneous feature fetching costs from different servers, achieving an O(log 3 k) competitive ratio in terms of overall feature-fetching communication cost (where k is the cache capacity). We evaluate practical performance of our caching design with testbed experiments, and show that our design achieves up to 5.4x convergence speed-up.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Eliminating Data Processing Bottlenecks in GNN Training over Large Graphs via Two-level Feature CompressionYuxin Ma, Ping Gong, Tianming Wu, Jiawei Yi et al.VLDB 2024 · 10 citations
- Accelerating Distributed Graph Learning by Using Collaborative In-Network Multicast and AggregationZhaoyi Li, Jiawei Huang, Yijun Li, Jingling Liu et al.USENIX ATC 2025 · 3 citations
Builds on8
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- P3: Distributed Deep Graph Learning at ScaleSwapnil Gandhi, Anand Padmanabha IyerOSDI 2021 · 192 citations
- Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless ThreadsJohn Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng et al.OSDI 2021 · 175 citations
- Random Walk Graph Neural NetworksGiannis Nikolentzos, Michalis VazirgiannisNeurIPS 2020 · 172 citations
- GNNLab: a factored system for sample-based GNN training over GPUsJianbang Yang, Dahai Tang, Xiaoniu Song, Lei Wang et al.EuroSys 2022 · 105 citations
Related papers
- BGL: GPU-Efficient GNN Training by Optimizing Graph Data I/O and PreprocessingTianfeng Liu, Yangrui Chen, Dan Li, Chuan Wu et al.NSDI 2023
- Expediting Distributed GNN Training with Feature-only Partition and Optimized Communication PlanningBingqian Du, Jun Liu, Ziyue Luo, Chuan Wu et al.INFOCOM 2024 · 4 citations
- On Pipelined GCN with Communication-Efficient Sampling and Inclusion-Aware CachingShulin Wang, Qiang Yu, Xiong Wang, Yuqing Li et al.INFOCOM 2024
- Optimizing Task Placement and Online Scheduling for Distributed GNN Training AccelerationZiyue Luo, Yixin Bao, Chuan WuINFOCOM 2022 · 12 citations
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal et al.SC 2021 · 35 citations
