CoGNN: Efficient Scheduling for Concurrent GNN Training on GPUs
Qingxiao Sun, Yi Liu, Hailong Yang, Ruizhe Zhang, Ming Dun, Mingzhen Li, Xiaoyan Liu, Wencong Xiao, Yong Li, Zhongzhi Luan, Depei Qian
摘要
Graph neural networks (GNNs) suffer from low GPU utilization due to frequent memory accesses. Existing concurrent training mechanisms cannot be directly adapted to GNNs because they fail to consider the impact of input irregularity. This requires pre-profiling the memory footprint of concurrent tasks based on input dimensions to ensure successful co-location on GPU. Moreover, massive training tasks generated from scenarios such as hyper-parameter tuning require flexible scheduling strategies. To address these problems, we propose CoGNN that enables efficient management of GNN training tasks on GPUs. Specifically, the CoGNN organizes the tasks in a queue and estimates the memory consumption of each task based on cost functions at operator basis. In addition, the CoGNN implements scheduling policies to generate task groups, which are iteratively submitted for execution. The experiment results show that the CoGNN can achieve shorter completion and queuing time for training tasks from diverse GNN models.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- PiPAD: Pipelined and Parallel Dynamic GNN Training on GPUsChunyang Wang, Desen Sun, Yuebin BaiPPoPP 2023 · 被引用 27 次
- Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU ClustersWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye 等EuroSys 2025 · 被引用 13 次
- TorchGT: A Holistic System for Large-Scale Graph Transformer TrainingMeng Zhang, Jie Sun, Qinghao Hu, Peng Sun 等SC 2024 · 被引用 7 次
- AutoGNN: End-to-End Hardware-Driven Graph Preprocessing for Enhanced GNN PerformanceSeungkwan Kang, Seungjun Lee, Donghyun Gouk, Miryeong Kwon 等HPCA 2026
- UniTG: A Unified System for Efficient and Seamless Textual Graph LearningMeng Zhang, Zhisheng Ye, Qiyu Liu, Jingshu Peng 等VLDB 2026
相关 Paper
- Optimizing Task Placement and Online Scheduling for Distributed GNN Training AccelerationZiyue Luo, Yixin Bao, Chuan WuINFOCOM 2022 · 被引用 12 次
- ElasGNN: An Elastic Training Framework for Distributed GNN TrainingSiqi Wang, Hailong Yang, Pengbo Wang, Hongliang Cao 等PPoPP 2026
- WiseGraph: Optimizing GNN with Joint Workload Partition of Graph and OperationsKezhao Huang, Jidong Zhai, Liyan Zheng, Haojie Wang 等EuroSys 2024 · 被引用 11 次
- NeutronOrch: Rethinking Sample-based GNN Training under CPU-GPU Heterogeneous EnvironmentsXin Ai, Qiange Wang, Chunyu Cao, Yanfeng Zhang 等VLDB 2024 · 被引用 19 次
- Adaptive Parallel Training for Graph Neural NetworksKaihao Ma, Renjie Liu, Xiao Yan, Zhenkun Cai 等PPoPP 2025
