Near-Optimal Wafer-Scale Reduce
Piotr Luczynski, Lukas Gianinazzi, Patrick Iff, Leighton Wilson, Daniele De Sensi, Torsten Hoefler
Abstract
Efficient Reduce and AllReduce communication collectives are a critical cornerstone of high-performance computing (HPC) applications. We present the first systematic investigation of Reduce and AllReduce on the Cerebras Wafer-Scale Engine (WSE). This architecture has been shown to achieve unprecedented performance both for machine learning workloads and other computational problems like FFT. We introduce a performance model to estimate the execution time of algorithms on the WSE and validate our predictions experimentally for a wide range of input sizes. In addition to existing implementations, we design and implement several new algorithms specifically tailored to the architecture. Moreover, we establish a lower bound for the runtime of a Reduce operation on the WSE. Based on our model, we automatically generate code that achieves near-optimal performance across the whole range of input sizes. Experiments demonstrate that our new Reduce and AllReduce algorithms outperform the current vendor solution by up to 3.27×. Additionally, our model predicts performance with less than 4% error. The proposed communication collectives increase the range of HPC applications that can benefit from the high throughput of the WSE. Our model-driven methodology demonstrates a disciplined approach that can lead the way to further algorithmic advancements on wafer-scale architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1bdf21c1-5649-4581-bef0-99c9bb566cf1Cited by top-tier papers2
- WaferLLM: Large Language Model Inference at Wafer ScaleCongjie He, Yeqi Huang, Pei Mu, Ziming Miao et al.OSDI 2025 · 20 citations
- Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale IntegrationYinxiao Feng, Kaisheng MaSC 2024 · 10 citations
Builds on6
- An in-depth analysis of the slingshot interconnectDaniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth et al.SC 2020 · 122 citations
- Fast stencil-code computation on a wafer-scale processorKamil Rocki, Dirk Van Essendelft, Ilya Sharapov, Robert Schreiber et al.SC 2020 · 69 citations
- Flare: flexible in-network allreduceDaniele De Sensi, Salvatore Di Girolamo, Saleh Ashkboos, Shigang Li et al.SC 2021 · 49 citations
- Swing: Short-cutting Rings for Higher Bandwidth AllreduceDaniele De Sensi, Tommaso Bonato, David Saam, Torsten HoeflerNSDI 2024 · 48 citations
- Accelerating bandwidth-bound deep learning inference with main-memory acceleratorsBenjamin Y. Cho, Jeageun Jung, Mattan ErezSC 2021 · 24 citations
Related papers
- An MLIR Lowering Pipeline for Stencils at Wafer-ScaleNicolai Stawinoga, David Katz, Anton Lydike, Justs Zarins et al.ASPLOS 2026
- Scalable Distributed High-Order Stencil ComputationsMathias Jacquelin, Mauricio Araya-Polo, Jie MengSC 2022 · 17 citations
- Automated Code Generation of High-Order Stencils for a Dataflow ArchitectureRyuichi Sai, John M. Mellor-Crummey, Jinfan Xu, Mauricio Araya-PoloSC 2024 · 8 citations
- TidalMesh: Topology-Driven AllReduce Collective Communication for Mesh TopologyDongkyun Lim, John KimHPCA 2025 · 12 citations
- ConBin: a Performance-Convergence Framework for Wafer-Scale Chip BinningHuiqing Xu, Mengdi Wang, Yinhe Han, Ying WangISCA 2026
