SiP-ML: high-bandwidth optical network interconnects for machine learning training
Mehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu, Madeleine Glick, Keren Bergman, Amin Vahdat, Benjamin Klenk, Eiman Ebrahimi
摘要
This paper proposes optical network interconnects as a key enabler for building high-bandwidth ML training clusters with strong scaling properties. Our design, called SiP-ML, accelerates the training time of popular DNN models using silicon photonics links capable of providing multiple terabits-per-second of bandwidth per GPU. SiP-ML partitions the training job across GPUs with hybrid data and model parallelism while ensuring the communication pattern can be supported efficiently on the network interconnect. We develop task partitioning and device placement methods that take the degree and reconfiguration latency of optical interconnects into account. Simulations using real DNN models show that, compared to the state-of-the-art electrical networks, our approach improves training time by 1.3--9.1x.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper16
- TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training JobsWeiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi 等NSDI 2023 · 被引用 215 次
- CASSINI: Network-Aware Job Scheduling in Machine Learning ClustersSudarsanan Rajasekaran, Manya Ghobadi, Aditya AkellaNSDI 2024 · 被引用 144 次
- In-Network Aggregation with Transport Transparency for Distributed TrainingShuo Liu, Qiaoling Wang, Junyi Zhang, Wenfei Wu 等ASPLOS 2023 · 被引用 46 次
- THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic CompressionMinghao Li, Ran Ben Basat, Shay Vargaftik, ChonLam Lao 等NSDI 2024 · 被引用 44 次
- Lightning: A Reconfigurable Photonic-Electronic SmartNIC for Fast and Energy-Efficient InferenceZhizhen Zhong, Mingran Yang, Jay Lang, Christian Williams 等SIGCOMM 2023 · 被引用 25 次
相关 Paper
- Opus: Photonic Rail-Optimized Fabric in ML DatacentersEric Ding, Barry Lyu, Bhaskar Kataria, Rachee SinghSIGCOMM 2026
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson 等NSDI 2021
- Mithril: A Scalable System for Deep GNN TrainingJingji Chen, Zhuoming Chen, Xuehai QianHPCA 2025 · 被引用 1 次
- Scaling Deep-Learning Inference with Chiplet-based Architecture and Photonic InterconnectsYuan Li, Ahmed Louri, Avinash KaranthDAC 2021 · 被引用 21 次
- TrainBox: An Extreme-Scale Neural Network Training Server Architecture by Systematically Balancing OperationsPyeongsu Park, Heetaek Jeong, Jangwoo KimMICRO 2020 · 被引用 11 次
