FRED: A Wafer-scale Fabric for 3D Parallel DNN Training
Saeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta, Tushar Krishna
Abstract
Wafer-scale systems are an emerging technology that tightly integrates high-end accelerator chiplets with high-speed wafer-scale interconnects, enabling low-latency and high-bandwidth connectivity. This makes them a promising platform for deep neural network (DNN) training. However, current network-on-wafer topologies, such as 2D Meshes, lack the flexibility needed to support various parallelization strategies effectively. In this paper, we propose Fred, a wafer-scale fabric architecture tailored to the unique communication needs of DNN training. Fred creates a distributed on-wafer topology with tiny microswitches, providing nonblocking connectivity for collective communications between arbitrary groups of accelerators and enabling in-switch collective support. Our results show that for sample parallelization strategies, Fred can improve the average end-to-end training time of ResNet-152, Transformer-17B, GPT-3, and Transformer-1T by 1.76×, 1.87×, 1.34×, and 1.4×, respectively, compared to a baseline wafer-scale Mesh.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext afc8eab2-9fc6-4f22-a87e-5a4233cbfef1Cited by top-tier papers3
- Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model InferenceYiqi Liu, Yudong Pan, Mengdi Wang, Shixin Zhao et al.ASPLOS 2026 · 1 citation
- Disdp: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel TrainingMo Sun, Zihan Yang, Changyue Liao, Yingtao Li et al.ISCA 2026 · 1 citation
- Patterns Behind Chaos: Forecasting Data Movement for Efficient Large-Scale Moe LLM InferenceZhongkai Yu, Yue Guan, Zihao Yu, Chenyang Zhou et al.ISCA 2026
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella et al.HPCA 2020 · 490 citations
Related papers
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 67 citations
- SuperMesh: Energy-Efficient Collective Communications for AcceleratorsSabuj Laskar, Pranati Majhi, Abdullah Muzahid, Eun Jung KimMICRO 2025 · 3 citations
- Scaling Graph Neural Network Training via Geometric OptimizationFangzhou Ye, Lingxiang Yin, Hao ZhengHPCA 2026
- TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training JobsWeiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi et al.NSDI 2023 · 215 citations
- PD Constraint-aware Physical/Logical Topology Co-Design for Network on WaferQize Yang, Taiquan Wei, Sihan Guan, Chengran Li et al.ISCA 2025 · 16 citations
