Resiliency at Scale: Managing Google's TPUv4 Machine Learning Supercomputer
Yazhou Zu, Alireza Ghaffarkhah, Hoang-Vu Dang, Brian Towles, Steven Hand, Safeen Huda, Adekunle Bello, Alexander Kolbasov, Arash Rezaei, Dayou Du, Steve Lacy, Hang Wang
摘要
TPUv4 (Tensor Processing Unit) is Google's 3rd generation accelerator for machine learning training, deployed as a 4096-node supercomputer with a custom 3D torus interconnect. In this paper, we describe our experience designing and operating the software infrastructure that allows TPUv4 supercomputers to operate at scale, including features for automatic fault resiliency and hardware recovery. We adopt a software-defined networking (SDN) approach to manage TPUv4's high-bandwidth inter-chip interconnect (ICI) fabric, using optical circuit switching to dynamically configure routes to work around machine, chip and link failures. Our infrastructure detects failures and automatically triggers reconfiguration to minimize disruption to running workloads, as well as initiating remediation and repair workflows for the affected components. Similar techniques interface with maintenance and upgrade workflows for both hardware and software. Our dynamic reconfiguration approach allows our TPUv4 supercomputers to achieve 99.98% system availability, gracefully handling hardware outages experienced by 1% of the training jobs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts TrainingXudong Liao, Yijun Sun, Han Tian, Xinchen Wan 等SIGCOMM 2025 · 被引用 14 次
- ReCycle: Resilient Training of Large DNNs using Pipeline AdaptationSwapnil Gandhi, Mark Zhao, Athinagoras Skiadopoulos, Christos KozyrakisSOSP 2024 · 被引用 13 次
- Reducing Energy Bloat in Large Model TrainingJae-Won Chung, Yile Gu, Insu Jang, Luoxi Meng 等SOSP 2024 · 被引用 12 次
- ResCCL: Resource-Efficient Scheduling for Collective CommunicationTongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao 等SIGCOMM 2025 · 被引用 11 次
- SkeletonHunter: Diagnosing and Localizing Network Failures in Containerized Large Model TrainingWei Liu, Kun Qian, Zhenhua Li, Tianyin Xu 等SIGCOMM 2025 · 被引用 8 次
它引用的顶会 Paper4
- Jupiter evolving: transforming google's datacenter network via optical circuit switches and software-defined networkingLeon Poutievski, Omid Mashayekhi, Joon Ong, Arjun Singh 等SIGCOMM 2022 · 被引用 230 次
- TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training JobsWeiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi 等NSDI 2023 · 被引用 215 次
- Orion: Google's Software-Defined Networking Control PlaneAndrew D. Ferguson, Steve D. Gribble, Chi-Yao Hong, Charles Killian 等NSDI 2021 · 被引用 95 次
- Experiences with Modeling Network Topologies at Multiple Levels of AbstractionJeffrey C. Mogul, Drago Goricanec, Martin Pool, Anees Shaikh 等NSDI 2020 · 被引用 43 次
相关 Paper
- Lightwave Fabrics: At-Scale Optical Circuit Switching for Datacenter and Machine Learning SystemsHong Liu, Ryohei Urata, Kevin Yasumura, Xiang Zhou 等SIGCOMM 2023 · 被引用 63 次
- A software-defined tensor streaming multiprocessor for large-scale machine learningDennis Abts, Garrin Kimmell, Andrew C. Ling, John Kim 等ISCA 2022 · 被引用 46 次
- Understanding and Mitigating Hardware Failures in Deep Learning Training SystemsYi He, Mike Hutton, Steven Chan, Robert De Gruijl 等ISCA 2023 · 被引用 52 次
- Reconfigurable Torus Fabrics for Multi-tenant MLAbhishek Vijaya Kumar, Eric Ding, Arjun Devraj, Darius Bunandar 等ASPLOS 2026 · 被引用 1 次
- SiP-ML: high-bandwidth optical network interconnects for machine learning trainingMehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu 等SIGCOMM 2021 · 被引用 94 次
