An MLIR Lowering Pipeline for Stencils at Wafer-Scale
Nicolai Stawinoga, David Katz, Anton Lydike, Justs Zarins, Nick Brown, George Bisbas, Tobias Grosser
Abstract
The Cerebras Wafer-Scale Engine (WSE) delivers performance at an unprecedented scale of over 900,000 compute units, all connected via a single-wafer on-chip interconnect. Initially designed for AI, the WSE architecture is also well-suited for High Performance Computing (HPC). However, its distributed asynchronous programming model diverges significantly from the simple sequential or bulk-synchronous programs that one would typically derive for a given mathematical program description. Targeting the WSE requires a bespoke re-implementation when porting existing code. The absence of WSE support in compilers such as MLIR, meant that there was little hope for automating this process.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd0eaf56-568b-42bd-83db-6231cb75b6ccBuilds on6
- Fast stencil-code computation on a wafer-scale processorKamil Rocki, Dirk Van Essendelft, Ilya Sharapov, Robert Schreiber et al.SC 2020 · 69 citations
- Scalable Distributed High-Order Stencil ComputationsMathias Jacquelin, Mauricio Araya-Polo, Jie MengSC 2022 · 17 citations
- IRDL: an IR definition language for SSA compilersMathieu Fehr, Jeff Niu, River Riddle, Mehdi Amini et al.PLDI 2022 · 13 citations
- A shared compilation stack for distributed-memory parallelism in stencil DSLsGeorge Bisbas, Anton Lydike, Emilien Bauer, Nick Brown et al.ASPLOS 2024 · 11 citations
- Automated Code Generation of High-Order Stencils for a Dataflow ArchitectureRyuichi Sai, John M. Mellor-Crummey, Jinfan Xu, Mauricio Araya-PoloSC 2024 · 8 citations
Related papers
- Near-Optimal Wafer-Scale ReducePiotr Luczynski, Lukas Gianinazzi, Patrick Iff, Leighton Wilson et al.HPDC 2024 · 7 citations
- Wavel: A Fast and Efficient Compilation System for Wafer-Scale AcceleratorsYeqi Huang, Congjie He, Haocheng Xiao, Yanwei Ye et al.SOSP 2026
- WaferLLM: Large Language Model Inference at Wafer ScaleCongjie He, Yeqi Huang, Pei Mu, Ziming Miao et al.OSDI 2025 · 20 citations
- ConBin: a Performance-Convergence Framework for Wafer-Scale Chip BinningHuiqing Xu, Mengdi Wang, Yinhe Han, Ying WangISCA 2026
- ScaleHLS: A New Scalable High-Level Synthesis Framework on Multi-Level Intermediate RepresentationHanchen Ye, Cong Hao, Jianyi Cheng, Hyunmin Jeong et al.HPCA 2022 · 77 citations
