Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale Integration
Yinxiao Feng, Kaisheng Ma
摘要
Existing high-performance computing (HPC) interconnection architectures are based on high-radix switches, which limits the injection/local performance and introduces latency/energy/cost overhead. The new wafer-scale packaging and high-speed wireline technologies provide high-density, low-latency, and high-bandwidth connectivity, thus promising to support direct-connected high-radix interconnection architecture.In this paper, we propose a wafer-based interconnection architecture called Switch-Less-Dragonfly-on-Wafers. By utilizing distributed high-bandwidth networks-on-chip-on-wafer, costly high-radix switches of the Dragonfly topology are eliminated while increasing the injection/local throughput and maintaining the global throughput. Based on the proposed architecture, we also introduce baseline and improved deadlock-free minimal/non-minimal routing algorithms with only one additional virtual channel. Extensive evaluations show that the Switch-Less-Dragonfly-on-Wafers outperforms the traditional switch-based Dragonfly in both cost and performance. Similar approaches can be applied to other switch-based direct topologies, thus promising to power future large-scale supercomputers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale ChipsHuizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang 等HPCA 2026 · 被引用 2 次
- WATOS: Efficient LLM Training Strategies and Architecture Co-Exploration for Wafer-Scale ChipHuizheng Wang, Zichuan Wang, Hongbin Wang, Jingxiang Hou 等HPCA 2026 · 被引用 2 次
- Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model InferenceYiqi Liu, Yudong Pan, Mengdi Wang, Shixin Zhao 等ASPLOS 2026 · 被引用 1 次
它引用的顶会 Paper10
- An in-depth analysis of the slingshot interconnectDaniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth 等SC 2020 · 被引用 122 次
- Designing a 2048-Chiplet, 14336-Core Waferscale ProcessorSaptadeep Pal, Jingyang Liu, Irina Alam, Nicholas Cebry 等DAC 2021 · 被引用 59 次
- Swing: Short-cutting Rings for Higher Bandwidth AllreduceDaniele De Sensi, Tommaso Bonato, David Saam, Torsten HoeflerNSDI 2024 · 被引用 48 次
- HexaMesh: Scaling to Hundreds of Chiplets with an Optimized Chiplet ArrangementPatrick Iff, Maciej Besta, Matheus A. Cavalcante, Tim Fischer 等DAC 2023 · 被引用 22 次
- PolarFly: A Cost-Effective and Flexible Low-Diameter TopologyKartik Lakhotia, Maciej Besta, Laura Monroe, Kelly Isham 等SC 2022 · 被引用 21 次
相关 Paper
- Ring Road: A Scalable Polar-Coordinate-based 2D Network-on-Chip ArchitectureYinxiao Feng, Wei Li, Kaisheng MaMICRO 2024 · 被引用 2 次
- Waferscale Network SwitchesShuangliang Chen, Saptadeep Pal, Rakesh KumarISCA 2024 · 被引用 18 次
- WaferBRAIN: Whole-Brain Scale Neuromorphic Architecture Based on Wafer-Scale IntegrationYukun Feng, Hao Jia, Liangyu Gan, Haoming Chu 等ISCA 2026
- A Scalable Methodology for Designing Efficient Interconnection Network of ChipletsYinxiao Feng, Dong Xiang, Kaisheng MaHPCA 2023 · 被引用 38 次
- FRED: A Wafer-scale Fabric for 3D Parallel DNN TrainingSaeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta 等ISCA 2025 · 被引用 8 次
