Near-Optimal Wafer-Scale Reduce
Piotr Luczynski, Lukas Gianinazzi, Patrick Iff, Leighton Wilson, Daniele De Sensi, Torsten Hoefler
摘要
Efficient Reduce and AllReduce communication collectives are a critical cornerstone of high-performance computing (HPC) applications. We present the first systematic investigation of Reduce and AllReduce on the Cerebras Wafer-Scale Engine (WSE). This architecture has been shown to achieve unprecedented performance both for machine learning workloads and other computational problems like FFT. We introduce a performance model to estimate the execution time of algorithms on the WSE and validate our predictions experimentally for a wide range of input sizes. In addition to existing implementations, we design and implement several new algorithms specifically tailored to the architecture. Moreover, we establish a lower bound for the runtime of a Reduce operation on the WSE. Based on our model, we automatically generate code that achieves near-optimal performance across the whole range of input sizes. Experiments demonstrate that our new Reduce and AllReduce algorithms outperform the current vendor solution by up to 3.27×. Additionally, our model predicts performance with less than 4% error. The proposed communication collectives increase the range of HPC applications that can benefit from the high throughput of the WSE. Our model-driven methodology demonstrates a disciplined approach that can lead the way to further algorithmic advancements on wafer-scale architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- WaferLLM: Large Language Model Inference at Wafer ScaleCongjie He, Yeqi Huang, Pei Mu, Ziming Miao 等OSDI 2025 · 被引用 20 次
- Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale IntegrationYinxiao Feng, Kaisheng MaSC 2024 · 被引用 10 次
它引用的顶会 Paper6
- An in-depth analysis of the slingshot interconnectDaniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth 等SC 2020 · 被引用 122 次
- Fast stencil-code computation on a wafer-scale processorKamil Rocki, Dirk Van Essendelft, Ilya Sharapov, Robert Schreiber 等SC 2020 · 被引用 69 次
- Flare: flexible in-network allreduceDaniele De Sensi, Salvatore Di Girolamo, Saleh Ashkboos, Shigang Li 等SC 2021 · 被引用 49 次
- Swing: Short-cutting Rings for Higher Bandwidth AllreduceDaniele De Sensi, Tommaso Bonato, David Saam, Torsten HoeflerNSDI 2024 · 被引用 48 次
- Accelerating bandwidth-bound deep learning inference with main-memory acceleratorsBenjamin Y. Cho, Jeageun Jung, Mattan ErezSC 2021 · 被引用 24 次
相关 Paper
- An MLIR Lowering Pipeline for Stencils at Wafer-ScaleNicolai Stawinoga, David Katz, Anton Lydike, Justs Zarins 等ASPLOS 2026
- Scalable Distributed High-Order Stencil ComputationsMathias Jacquelin, Mauricio Araya-Polo, Jie MengSC 2022 · 被引用 17 次
- Automated Code Generation of High-Order Stencils for a Dataflow ArchitectureRyuichi Sai, John M. Mellor-Crummey, Jinfan Xu, Mauricio Araya-PoloSC 2024 · 被引用 8 次
- TidalMesh: Topology-Driven AllReduce Collective Communication for Mesh TopologyDongkyun Lim, John KimHPCA 2025 · 被引用 12 次
- ConBin: a Performance-Convergence Framework for Wafer-Scale Chip BinningHuiqing Xu, Mengdi Wang, Yinhe Han, Ying WangISCA 2026
