SC2024Top-tier venue
Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale Integration
Yinxiao Feng, Kaisheng Ma
Abstract
Existing high-performance computing (HPC) interconnection architectures are based on high-radix switches, which limits the injection/local performance and introduces latency/energy/cost overhead. The new wafer-scale packaging and high-speed wireline technologies provide high-density, low-latency, and high-bandwidth connectivity, thus promising to support direct-connected high-radix interconnection architecture.In this paper, we propose a wafer-based interconnection architecture called Switch-Less-Dragonfly-on-Wafers. By utilizing distributed high-bandwidth networks-on-chip-on-wafer, costly high-radix switches of the Dragonfly topology are eliminated while increasing the injection/local throughput and maintaining the global throughput. Based on the proposed architecture, we also introduce baseline and improved deadlock-free minimal/non-minimal routing algorithms with only one additional virtual channel. Extensive evaluations show that the Switch-Less-Dragonfly-on-Wafers outperforms the traditional switch-based Dragonfly in both cost and performance. Similar approaches can be applied to other switch-based direct topologies, thus promising to power future large-scale supercomputers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e2a19bf-a1c4-4b81-ae87-af2f1bcdcf0aCited by top-tier papers3
- TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale ChipsHuizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang et al.HPCA 2026 · 2 citations
- WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale ChipHuizheng Wang, Zichuan Wang, Hongbin Wang, Jingxiang Hou et al.HPCA 2026 · 2 citations
- Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model InferenceYiqi Liu, Yudong Pan, Mengdi Wang, Shixin Zhao et al.ASPLOS 2026 · 1 citation
Builds on10
- An in-depth analysis of the slingshot interconnectDaniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth et al.SC 2020 · 122 citations
- Designing a 2048-Chiplet, 14336-Core Waferscale ProcessorSaptadeep Pal, Jingyang Liu, Irina Alam, Nicholas Cebry et al.DAC 2021 · 59 citations
- Swing: Short-cutting Rings for Higher Bandwidth AllreduceDaniele De Sensi, Tommaso Bonato, David Saam, Torsten HoeflerNSDI 2024 · 48 citations
- HexaMesh: Scaling to Hundreds of Chiplets with an Optimized Chiplet ArrangementPatrick Iff, Maciej Besta, Matheus A. Cavalcante, Tim Fischer et al.DAC 2023 · 22 citations
- PolarFly: A Cost-Effective and Flexible Low-Diameter TopologyKartik Lakhotia, Maciej Besta, Laura Monroe, Kelly Isham et al.SC 2022 · 21 citations
Related papers
- Ring Road: A Scalable Polar-Coordinate-based 2D Network-on-Chip ArchitectureYinxiao Feng, Wei Li, Kaisheng MaMICRO 2024 · 2 citations
- Waferscale Network SwitchesShuangliang Chen, Saptadeep Pal, Rakesh KumarISCA 2024 · 18 citations
- WaferBRAIN: Whole-Brain Scale Neuromorphic Architecture Based on Wafer-Scale IntegrationYukun Feng, Hao Jia, Liangyu Gan, Haoming Chu et al.ISCA 2026
- A Scalable Methodology for Designing Efficient Interconnection Network of ChipletsYinxiao Feng, Dong Xiang, Kaisheng MaHPCA 2023 · 38 citations
- FRED: A Wafer-scale Fabric for 3D Parallel DNN TrainingSaeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta et al.ISCA 2025 · 8 citations
