DUCATI: A Dual-Cache Training System for Graph Neural Networks on Giant Graphs with the GPU
Xin Zhang, Yanyan Shen, Yingxia Shao, Lei Chen
Abstract
Recently Graph Neural Networks (GNNs) have achieved great success in many applications. The mini-batch training has become the de-facto way to train GNNs on giant graphs. However, the mini-batch generation task is extremely expensive which slows down the whole training process. Researchers have proposed several solutions to accelerate the mini-batch generation, however, they (1) fail to exploit the locality of the adjacency matrix, (2) cannot fully utilize the GPU memory, and (3) suffer from the poor adaptability to diverse workloads. In this work, we propose DUCATI, aDual-Cache system to overcome these drawbacks. In addition to the traditionalNfeat-Cache, DUCATI introduces a newAdj-Cache to further accelerate the mini-batch generation and better utilize GPU memory. DUCATI develops a workload-awareDual-Cache Allocator which adaptively finds the best cache allocation plan under different settings. We compare DUCATI with various GNN training systems on four billion-scale graphs under diverse workload settings. The experimental results show that in terms of training time, DUCATI can achieve up to 3.33 times speedup (2.07 times on average) compared to DGL and up to 1.54 times speedup (1.32 times on average) compared to the state-of-the-artSingle-Cache systems. We also analyze the time-accuracy trade-offs of DUCATI and four state-of-the-art GNN training systems. The analysis results offer users some guidelines on system selection regarding different input sizes and hardware resources.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get dfff01b0-e2ec-4710-a3e9-d124afb11917Cited by top-tier papers15
- ETC: Efficient Training of Temporal Graph Neural Networks over Large-scale Dynamic GraphsShihong Gao, Yiming Li, Yanyan Shen, Yingxia Shao et al.VLDB 2024 · 32 citations
- Graph Neural Network Training Systems: A Performance Comparison of Full-Graph and Mini-BatchSaurabh Bajaj, Hui Guan, Marco Serafini, Juelin Liu et al.VLDB 2025 · 19 citations
- NeutronOrch: Rethinking Sample-based GNN Training under CPU-GPU Heterogeneous EnvironmentsXin Ai, Qiange Wang, Chunyu Cao, Yanfeng Zhang et al.VLDB 2024 · 19 citations
- DAHA: Accelerating GNN Training with Data and Hardware Aware Execution PlanningZhiyuan Li, Xun Jian, Yue Wang, Yingxia Shao et al.VLDB 2024 · 18 citations
- Eliminating Data Processing Bottlenecks in GNN Training over Large Graphs via Two-level Feature CompressionYuxin Ma, Ping Gong, Tianming Wu, Jiawei Yi et al.VLDB 2024 · 10 citations
Related papers
- Efficient GNN Training on Giant Graphs with Collective Batching and SchedulingXin Zhang, Yanyan Shen, Yingxia Shao, Haoyang Li et al.VLDB 2026
- TAC: Cache-Based System for Accelerating Billion-Scale GNN Training on Multi-GPU PlatformZhiqiang Liang, Hongyu Gao, Jue Wang, Fang Liu et al.PPoPP 2026
- BGL: GPU-Efficient GNN Training by Optimizing Graph Data I/O and PreprocessingTianfeng Liu, Yangrui Chen, Dan Li, Chuan Wu et al.NSDI 2023
- Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN TrainingJie Sun, Li Su, Zuocheng Shi, Wenting Shen et al.USENIX ATC 2023 · 5 citations
- MariusGNN: Resource-Efficient Out-of-Core Training of Graph Neural NetworksRoger Waleffe, Jason Mohoney, Theodoros Rekatsinas, Shivaram VenkataramanEuroSys 2023 · 40 citations
