Exploring GPU-to-GPU Communication: Insights into Supercomputer Interconnects
Daniele De Sensi, Lorenzo Pichetti, Flavio Vella, Tiziano De Matteis, Zebin Ren, Luigi Fusco, Matteo Turisini, Daniele Cesarini, Kurt Lust, Animesh Trivedi, Duncan Roweth, Filippo Spiga
摘要
Multi-GPU nodes are increasingly common in the rapidly evolving landscape of exascale supercomputers. On these systems, GPUs on the same node are connected through dedicated networks, with bandwidths up to a few terabits per second. However, gauging performance expectations and maximizing system efficiency is challenging due to different technologies, design options, and software layers. This paper comprehensively characterizes three supercomputers — Alps, Leonardo, and LUMI — each with a unique architecture and design. We focus on performance evaluation of intra-node and inter-node interconnects on up to 4,096 GPUs, using a mix of intra-node and inter-node benchmarks. By analyzing its limitations and opportunities, we aim to offer practical guidance to researchers, system architects, and software developers dealing with multi-GPU supercomputing. Our results show that there is untapped bandwidth, and there are still many opportunities for optimization, ranging from network to software optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed StorageSiyuan Shen, Tommaso Bonato, Zhiyi Hu, Pasquale Jordan 等SC 2025 · 被引用 7 次
- Bine Trees: Enhancing Collective Operations by Optimizing Communication LocalityDaniele De Sensi, Saverio Pasqualoni, Lorenzo Piarulli, Tommaso Bonato 等SC 2025 · 被引用 4 次
它引用的顶会 Paper3
- An in-depth analysis of the slingshot interconnectDaniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth 等SC 2020 · 被引用 122 次
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl 等SIGMOD 2020 · 被引用 99 次
- Frontier: Exploring ExascaleScott Atchley, Christopher Zimmer, John Lange, David E. Bernholdt 等SC 2023 · 被引用 91 次
相关 Paper
- Compass: Dissecting Communication and Computation Operators for Efficient LLM TrainingGuangyu Xiang, Lin Zhang, Haoxuan Yu, Xinglin Pan 等INFOCOM 2026
- MG-Join: A Scalable Join for Massively Parallel Multi-GPU ArchitecturesJohns Paul, Shengliang Lu, Bingsheng He, Chiew Tong LauSIGMOD 2021 · 被引用 31 次
- Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal PerspectiveSeokjin Go, Joongun Park, Spandan More, Hanjiang Wu 等MICRO 2025 · 被引用 7 次
- Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUsGuoliang He, Youhe Jiang, Wencong Xiao, Kaihua Jiang 等NeurIPS 2025 · 被引用 10 次
- Evaluating Multi-GPU Sorting with Modern InterconnectsTobias Maltenberger, Ivan Ilic, Ilin Tolovski, Tilmann RablSIGMOD 2022 · 被引用 24 次
