Scaling Graph Neural Network Training via Geometric Optimization
Fangzhou Ye, Lingxiang Yin, Hao Zheng
Abstract
Wafer-scale computing has emerged as an alternative solution to sustain performance scaling in the post-Moore era, driven by recent technology advancements such as chiplet integration. This enables considerable computing and storage capabilities on a single chip, making it capable of accommodating large machine learning models and datasets. Recent efforts have heralded the promise of wafer-scale architectures for deep learning inference and training. However, scaling the training of Graph Neural Networks in wafer-scale architecture remains a challenge and is relatively unexplored due to irregularities in gradient propagation as well as physical constraints from flat on-chip topologies. In this paper, we propose Aster, a topology-aware framework designed to efficiently support GNN training on arbitrary wafer-scale architectures. The proposed framework, as opposed to the current application or topology-specific heuristics, can be generalized to support any network topology and irregular GNN datasets. Specifically, we mathematically formulate commonly-seen network topologies in their geometric representation and prioritize communication efficiency during GNN workload partitioning and mapping. Based on the geometric representation, we propose a quadratic assignment problem solver to efficiently map irregular dataflows to a flat topology with reduced communication distance. The simulation results show that Aster can achieve performance speedup by, andin Mesh and speedup by, andin Torus on average compared to Mini-cut [1], ScalaGraph [2], ChunkV [3], and Chunk-E [4], respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0795b521-677e-46a0-9b1b-bdb9c3d63679Related papers
- FRED: A Wafer-scale Fabric for 3D Parallel DNN TrainingSaeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta et al.ISCA 2025 · 8 citations
- ElasGNN: An Elastic Training Framework for Distributed GNN TrainingSiqi Wang, Hailong Yang, Pengbo Wang, Hongliang Cao et al.PPoPP 2026
- NeutronHeter: Optimizing Distributed Graph Neural Network Training for Heterogeneous ClustersChunyu Cao, Xin Ai, Qiange Wang, Yanfeng Zhang et al.SIGMOD 2026 · 3 citations
- WiseGraph: Optimizing GNN with Joint Workload Partition of Graph and OperationsKezhao Huang, Jidong Zhai, Liyan Zheng, Haojie Wang et al.EuroSys 2024 · 11 citations
- NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor ParallelismXin Ai, Hao Yuan, Zeyu Ling, Qiange Wang et al.VLDB 2025 · 8 citations
