Zettafly: A Network Topology with Flexible Non-blocking Regions for Large-scale AI and HPC Systems
Dezun Dong, Ziyu Wang, Fei Lei
Abstract
Interconnection networks are playing an increasingly crucial role in achieving scalability and throughput for post-exascale and zettascale computing systems.Resource consumption characteristics of AI and HPC workloads are essential to designing effective network infrastructure.The operation practice of production supercomputing systems reveals that small and medium-sized jobs consume most compute-core hours, generating traffic that mainly utilizes a size-restricted portion of the network.Unfortunately, current topologies lack sufficient support for flexible partitioning of distinct concurrent jobs.This study bridges this gap by exploring a new tradeoff among scalability, throughput, and non-blocking regions.We present Zettafly, a family of low-diameter topologies with large-scale non-blocking sub-networks, allowing the majority of jobs to be isolated within a single sub-network.Zettafly is a two-layer structure, consisting of non-blocking groups and global routers, such that each minimally routed packet between groups traverses at most one global router.We also introduce simple and efficient adaptive routing algorithms for Zettafly.We conduct extensive simulations and analysis to evaluate the performance and cost of Zettafly against state-of-the-art topologies.The results show that Zettafly achieves great performance under typical HPC workloads, offers cost-effectiveness under multi-task mixed traffic, supports flexible configurations and incremental deployments, and is costeffective to deploy for post-exascale and zettascale systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 080eaa4e-336b-47f1-a9b5-e8b1d7c4a1d6Related papers
- A High-Performance Design, Implementation, Deployment, and Evaluation of The Slim Fly NetworkNils Blach, Maciej Besta, Daniele De Sensi, Jens Domke et al.NSDI 2024 · 13 citations
- An Evaluation of the Effect of Network Cost Optimization for Leadership Class SupercomputersAwais Khan, John R. Lange, Nick Hagerty, Edwin F. Posada et al.SC 2024 · 4 citations
- Study of Workload Interference with Intelligent Routing on DragonflyYao Kang, Xin Wang, Zhiling LanSC 2022 · 7 citations
- Architecture and performance studies of 3D-Hyper-FleX-LION for reconfigurable all-to-all HPC networksGengchen Liu, Roberto Proietti, Marjan Fariborz, Pouya Fotouhi et al.SC 2020 · 21 citations
- An in-depth analysis of the slingshot interconnectDaniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth et al.SC 2020 · 122 citations
