Reconfigurable Torus Fabrics for Multi-tenant ML
Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj, Darius Bunandar, Rachee Singh
摘要
We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML datacenters with Morphlux can improve the bandwidth of tenant compute allocations by up to 66%, reduce compute fragmentation by up to 70%, and minimize the blast radius of accelerator failures. We develop a novel end-to-end hardware prototype of Morphlux to demonstrate these performance benefits which translate to 1.72x improvement in finetuning throughput of ML models. By rapidly programming the server-scale fabric in our hardware testbed, Morphlux can replace a failed accelerator with a healthy one in 1.2 seconds.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Lightwave Fabrics: At-Scale Optical Circuit Switching for Datacenter and Machine Learning SystemsHong Liu, Ryohei Urata, Kevin Yasumura, Xiang Zhou 等SIGCOMM 2023 · 被引用 63 次
- SiP-ML: high-bandwidth optical network interconnects for machine learning trainingMehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu 等SIGCOMM 2021 · 被引用 94 次
- FRED: A Wafer-scale Fabric for 3D Parallel DNN TrainingSaeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta 等ISCA 2025 · 被引用 8 次
- Resiliency at Scale: Managing Google's TPUv4 Machine Learning SupercomputerYazhou Zu, Alireza Ghaffarkhah, Hoang-Vu Dang, Brian Towles 等NSDI 2024 · 被引用 46 次
- Lynx: A SmartNIC-driven Accelerator-centric Architecture for Network ServersMaroun Tork, Lina Maudlej, Mark SilbersteinASPLOS 2020 · 被引用 64 次
