Waferscale Network Switches
Shuangliang Chen, Saptadeep Pal, Rakesh Kumar
Abstract
In spite of being a key determinant of latency, cost, power, space, and capability of modern computer systems, network switch radix has not seen much growth over the years due to poor scaling of off-chip IO pitches and switch die sizes. We consider waferscale integration (WSI) as a way to increase the size of the switch substrate to be much bigger than a single die and ask the question: can we use WSI to enable network switches that have dramatically higher radix than today’s switches? We show that while a waferscale network switch can support up to 32x higher radix than state-of-the-art network switches when only area constraints are considered, the actual radix of a waferscale network switch is not area-limited. Rather, it is limited by a combination of internal bandwidth, external bandwidth, and power density. In fact, without optimizations, benefits of a waferscale network switch are minimal. To address the scalability bottlenecks, we propose a heterogeneous network switch design that reduces switch power by 30.8%-33.5% which, in turn, allows an increase in radix (by up to 4x) by increasing internal I/O bandwidth at the expense of energy efficiency. We also propose subswitch deradixing that increases the overall radix by 2x by decreasing the radix of the subswitches to alleviate the internal I/O bottleneck. We use Area I/O and Optical I/O schemes to alleviate the external I/O bandwidth bottlenecks of conventional SerDes-based external connectivity. In addition to scalability optimization, we present optimizations such as low latency buffering and proprietary routing that improve the performance of waferscale switches. Finally, we present a system architecture for a waferscale network switch that supports its port count, power delivery, and cooling requirements in a compact form factor. We show that the switch can be used to enable new computing systems such as single-switch datacenters and massive-scale singular GPUs. It can also lead to a dramatic reduction in datacenter network costs. Overall, this is the first work quantifying the benefits of waferscale switches and identifying and addressing the unique challenges and opportunities in building them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 431cfac0-e24a-4b99-9121-3d899fef452dCited by top-tier papers6
- Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale IntegrationYinxiao Feng, Kaisheng MaSC 2024 · 10 citations
- TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale ChipsHuizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang et al.HPCA 2026 · 2 citations
- WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale ChipHuizheng Wang, Zichuan Wang, Hongbin Wang, Jingxiang Hou et al.HPCA 2026 · 2 citations
- HDPAT: Hierarchical Distributed Page Address Translation for Wafer-Scale GPUsDaoxuan Xu, Ying Li, Yuwei Sun, Jie Ren et al.HPCA 2026 · 1 citation
- NetSparse: In-Network Acceleration of Distributed Sparse KernelsGerasimos Gerogiannis, Dimitrios Merkouriadis, Charles Block, Annus Zulfiqar et al.MICRO 2025 · 1 citation
Builds on2
- Jupiter evolving: transforming google's datacenter network via optical circuit switches and software-defined networkingLeon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh et al.SIGCOMM 2022 · 230 citations
- Aquila: A unified, low-latency fabric for datacenter networksDan Gibson, Hema Hariharan, Eric Lance, Moray McLaren et al.NSDI 2022 · 60 citations
Related papers
- Designing a 2048-Chiplet, 14336-Core Waferscale ProcessorSaptadeep Pal, Jingyang Liu, Irina Alam, Nicholas Cebry et al.DAC 2021 · 59 citations
- Sirius: A Flat Datacenter Network with Nanosecond Optical SwitchingHitesh Ballani, Paolo Costa, Raphael Behrendt, Daniel Cletheroe et al.SIGCOMM 2020 · 204 citations
- Evaluating Ruche Networks: Physically Scalable, Cost-Effective, Bandwidth-Flexible NoCsDai Cheol Jung, Michael B. TaylorISCA 2025 · 3 citations
- ACTINA: Adapting Circuit-Switching Techniques for AI Networking ArchitecturesZhenguo Wu, Benjamin Klenk, Larry Dennison, Keren BergmanSC 2025 · 4 citations
- FRED: A Wafer-scale Fabric for 3D Parallel DNN TrainingSaeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta et al.ISCA 2025 · 8 citations
