Hyperion: Co-Optimizing SSD Access and GPU Computation for Cost-Efficient GNN Training
Jie Sun, Mo Sun, Zheng Zhang, Zuocheng Shi, Jun Xie, Zihan Yang, Jie Zhang, Zeke Wang, Fei Wu
摘要
SSDs are traditionally regarded as a cheap but slow way to scale up GNN training. Several GNN systems explore cheap single-machine single-GPU out-of-core training but fall short in terms of TPC (throughput per monetary cost). The underlying reason is that the existing systems 1) overly focus on minimizing the number of SSD accesses, which results in substantial unnecessary overhead on the CPU side, or 2) exhaust all GPU parallelism to saturate SSD but fail to overlap SSD accesses with GNN computation. In this work, we present Hyperion, a cost-efficient system for terabyte-scale GNN training. We argue that co-optimizing GPU-initiated asynchronous SSD access and GNN computation pipeline enables us to only add cheap NVMe SSDs, rather than expensive GPU servers, to achieve in- memory-like throughput and thus maximal TPC of GNN training. However, this is non-trivial due to imbalanced workloads and interference among IO submission, IO completion, and cache lookup. To tackle the challenges, Hyperion proposes three key designs. First, Hyperion proposes the first GPU-initiated pipeline- friendly asynchronous disk IO stack, which only requires about 1% GPU cores to saturate SSD throughput and wastes no GPU cores between IO submission and completion to fully overlap disk IO and computation. Second, we propose a new GPU-managed, disaggregated, and unified cache that disaggregates cache lookup from disk IO and fully utilizes CPU/GPU memory hierarchy by a unified static cache policy. Third, we propose a GNN-aware general TPC-analytical model that precisely predicts TPC under diverse hardware settings and GNN models and provide a hint to guide users to select hardware, e.g., number of SSDs, under a limited budget to maximize TPC. Experiments demonstrate that Hyperion can improve the TPC by over 3.1x on terabyte-scale graphs compared to SOTA out-of-core baselines and improve 60 x TPC compared to distributed in-memory baselines.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Scalable Graph Indexing using GPUs for Approximate Nearest Neighbor SearchZhonggen Li, Xiangyu Ke, Yifan Zhu, Bocheng Yu 等SIGMOD 2026 · 被引用 5 次
- Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN TrainingZuocheng Shi, Jie Sun, Ziyu Song, Mo Sun 等SC 2025 · 被引用 3 次
- Bat: Efficient Generative Recommender Serving with Bipartite AttentionJie Sun, Shaohang Wang, Zimo Zhang, Zhengyu Liu 等ASPLOS 2026 · 被引用 1 次
- ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural NetworksPranjal Naman, Yogesh SimmhanHPDC 2026
相关 Paper
- MariusGNN: Resource-Efficient Out-of-Core Training of Graph Neural NetworksRoger Waleffe, Jason Mohoney, Theodoros Rekatsinas, Shivaram VenkataramanEuroSys 2023 · 被引用 40 次
- Ginex: SSD-enabled Billion-scale Graph Neural Network Training on a Single Machine via Provably Optimal In-memory CachingYeonhong Park, Sunhong Min, Jae W. LeeVLDB 2022 · 被引用 57 次
- Legion: Automatically Pushing the Envelope of Multi-GPU System for Billion-Scale GNN TrainingJie Sun, Li Su, Zuocheng Shi, Wenting Shen 等USENIX ATC 2023 · 被引用 5 次
- CaliEX: A Disk-Based Large-Scale GNN Training System with Joint Design of Caching and ExecutionCan Su, Haipeng Zhang, Hanyu Zhao, Wenting Shen 等ICDE 2025 · 被引用 1 次
- CoGNN: Efficient Scheduling for Concurrent GNN Training on GPUsQingxiao Sun, Yi Liu, Hailong Yang, Ruizhe Zhang 等SC 2022 · 被引用 11 次
