Two-level Graph Caching for Expediting Distributed GNN Training
Zhe Zhang, Ziyue Luo, Chuan Wu
摘要
Graph Neural Networks (GNNs) are increasingly popular due to excellent performance on learning graphstructured data in various domains. With fast expanding graph sizes and feature dimensions, distributed GNN training has been adopted, with multiple concurrent workers learning on different portions of a large graph. It has been observed that a main bottleneck in distributed GNN training lies in graph feature fetching across servers, which dominates the training time of each training iteration at each worker. This paper studies efficient feature caching on each worker to minimize feature fetching overhead, in order to expedite distributed GNN training. Current distributed GNN training systems largely adopt static caching of fixed neighbor nodes. We propose a novel two-level dynamic cache design exploiting both GPU memory and host memory at each worker, and design efficient two-level dynamic caching algorithms based on online optimization and a lookahead batching mechanism. Our dynamic caching algorithms consider node requesting probabilities and heterogeneous feature fetching costs from different servers, achieving an O(log 3 k) competitive ratio in terms of overall feature-fetching communication cost (where k is the cache capacity). We evaluate practical performance of our caching design with testbed experiments, and show that our design achieves up to 5.4x convergence speed-up.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Eliminating Data Processing Bottlenecks in GNN Training over Large Graphs via Two-level Feature CompressionYuxin Ma, Ping Gong, Tianming Wu, Jiawei Yi 等VLDB 2024 · 被引用 10 次
- Accelerating Distributed Graph Learning by Using Collaborative In-Network Multicast and AggregationZhaoyi Li, Jiawei Huang, Yijun Li, Jingling Liu 等USENIX ATC 2025 · 被引用 3 次
它引用的顶会 Paper8
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- P3: Distributed Deep Graph Learning at ScaleSwapnil Gandhi, Anand Padmanabha IyerOSDI 2021 · 被引用 192 次
- Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless ThreadsJohn Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng 等OSDI 2021 · 被引用 175 次
- Random Walk Graph Neural NetworksGiannis Nikolentzos, Michalis VazirgiannisNeurIPS 2020 · 被引用 172 次
- GNNLab: a factored system for sample-based GNN training over GPUsJianbang Yang, Dahai Tang, Xiaoniu Song, Lei Wang 等EuroSys 2022 · 被引用 105 次
相关 Paper
- BGL: GPU-Efficient GNN Training by Optimizing Graph Data I/O and PreprocessingTianfeng Liu, Yangrui Chen, Dan Li, Chuan Wu 等NSDI 2023
- Expediting Distributed GNN Training with Feature-only Partition and Optimized Communication PlanningBingqian Du, Jun Liu, Ziyue Luo, Chuan Wu 等INFOCOM 2024 · 被引用 4 次
- On Pipelined GCN with Communication-Efficient Sampling and Inclusion-Aware CachingShulin Wang, Qiang Yu, Xiong Wang, Yuqing Li 等INFOCOM 2024
- Optimizing Task Placement and Online Scheduling for Distributed GNN Training AccelerationZiyue Luo, Yixin Bao, Chuan WuINFOCOM 2022 · 被引用 12 次
- Efficient scaling of dynamic graph neural networksVenkatesan T. Chakaravarthy, Shivmaran S. Pandian, Saurabh Raje, Yogish Sabharwal 等SC 2021 · 被引用 35 次
