ConBin: a Performance-Convergence Framework for Wafer-Scale Chip Binning
Huiqing Xu, Mengdi Wang, Yinhe Han, Ying Wang
摘要
Wafer-scale chips (WSCs) like the Cerebras WaferScale Engines (WSEs) offer immense on-chip compute and bandwidth for large-scale Artificial Intelligence (AI) workloads, but their massive area makes defect-free fabrication impossible, posing severe yield and cost challenges. The traditional solution is chip binning, which grades chips by frequency or core count, but such metrics fail for WSCs as their performance is not only determined by core number but also heavily influenced by the communication irregularities caused by faults. These fault-induced variations cause large inter-chip performance divergence, forcing conservative binning thresholds that reduce premium-bin yield (i.e. the fraction of chips in the highest-performance bins) and the aggregate guaranteed performance delivered across bins (i.e., the total sellable effective compute capacity, SECC). Therefore, a new binning strategy and performance-convergence mechanisms are essential for the practical commercialization of WSCs. To address these challenges, we first propose Performance Binning, a novel paradigm that grades WSCs by their actual performance on target workloads rather than by core count or frequency. We further develop a Performance-Convergence Framework for WSC Binning (ConBIN) that unifies hardware design, software optimization, and performance binning to converge the inter-chip performance distribution and maximize binning yield. ConBIN employs automated, fault-correlationaware redundant interconnect design and post-silicon fault repair to reduce inter-chip divergence of core counts and topology structures. It then applies bin-aware workload mapping and finegrained communication scheduling guided by lightweight prebinning targets to suppress residual performance variance, which enables tighter chip bin thresholds and higher premium-bin yield, boosting overall SECC. Finally, ConBIN executes performance binning to determine binning thresholds that maximize total SECC. Evaluations show that ConBIN improves premium-bin yield by 2.80× and total SECC by 2.64× over state-of-the-art (SOTA) fault-tolerant methods on WSCs.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Near-Optimal Wafer-Scale ReducePiotr Luczynski, Lukas Gianinazzi, Patrick Iff, Leighton Wilson 等HPDC 2024 · 被引用 7 次
- An MLIR Lowering Pipeline for Stencils at Wafer-ScaleNicolai Stawinoga, David Katz, Anton Lydike, Justs Zarins 等ASPLOS 2026
- Si-Kintsugi: Towards Recovering Golden-Like Performance of Defective Many-Core Spatial Architectures for AIEdward Hanson, Shiyu Li, Guanglei Zhou, Feng Cheng 等MICRO 2023 · 被引用 3 次
- Wavel: A Fast and Efficient Compilation System for Wafer-Scale AcceleratorsYeqi Huang, Congjie He, Haocheng Xiao, Yanwei Ye 等SOSP 2026
- Designing a 2048-Chiplet, 14336-Core Waferscale ProcessorSaptadeep Pal, Jingyang Liu, Irina Alam, Nicholas Cebry 等DAC 2021 · 被引用 59 次
