SiP-ML: high-bandwidth optical network interconnects for machine learning training
Mehrdad Khani Shirkoohi, Manya Ghobadi, Mohammad Alizadeh, Ziyi Zhu, Madeleine Glick, Keren Bergman, Amin Vahdat, Benjamin Klenk, Eiman Ebrahimi
Abstract
This paper proposes optical network interconnects as a key enabler for building high-bandwidth ML training clusters with strong scaling properties. Our design, called SiP-ML, accelerates the training time of popular DNN models using silicon photonics links capable of providing multiple terabits-per-second of bandwidth per GPU. SiP-ML partitions the training job across GPUs with hybrid data and model parallelism while ensuring the communication pattern can be supported efficiently on the network interconnect. We develop task partitioning and device placement methods that take the degree and reconfiguration latency of optical interconnects into account. Simulations using real DNN models show that, compared to the state-of-the-art electrical networks, our approach improves training time by 1.3--9.1x.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 36240ac3-205b-4888-93ba-6c85f0fa34b0Cited by top-tier papers16
- TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training JobsWeiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi et al.NSDI 2023 · 215 citations
- CASSINI: Network-Aware Job Scheduling in Machine Learning ClustersSudarsanan Rajasekaran, Manya Ghobadi, Aditya AkellaNSDI 2024 · 144 citations
- In-Network Aggregation with Transport Transparency for Distributed TrainingShuo Liu, Qiaoling Wang, Junyi Zhang, Wenfei Wu et al.ASPLOS 2023 · 46 citations
- THC: Accelerating Distributed Deep Learning Using Tensor Homomorphic CompressionMinghao Li, Ran Ben Basat, Shay Vargaftik, ChonLam Lao et al.NSDI 2024 · 44 citations
- Lightning: A Reconfigurable Photonic-Electronic SmartNIC for Fast and Energy-Efficient InferenceZhizhen Zhong, Mingran Yang, Jay Lang, Christian Williams et al.SIGCOMM 2023 · 25 citations
Related papers
- Opus: Photonic Rail-Optimized Fabric in ML DatacentersEric Ding, Barry Lyu, Bhaskar Kataria, Rachee SinghSIGCOMM 2026
- Scaling Distributed Machine Learning with In-Network AggregationAmedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson et al.NSDI 2021
- Mithril: A Scalable System for Deep GNN TrainingJingji Chen, Zhuoming Chen, Xuehai QianHPCA 2025 · 1 citation
- Scaling Deep-Learning Inference with Chiplet-based Architecture and Photonic InterconnectsYuan Li, Ahmed Louri, Avinash KaranthDAC 2021 · 21 citations
- TrainBox: An Extreme-Scale Neural Network Training Server Architecture by Systematically Balancing OperationsPyeongsu Park, Heetaek Jeong, Jangwoo KimMICRO 2020 · 11 citations
