Lightwave Fabrics: At-Scale Optical Circuit Switching for Datacenter and Machine Learning Systems
Hong Liu, Ryohei Urata, Kevin Yasumura, Xiang Zhou, Roy Bannon, Jill Berger, Pedram Dashti, Norm Jouppi, Cedric F. Lam, Sheng Li, Erji Mao, Daniel Nelson
Abstract
We describe our experience developing what we believe to be the world's first large-scale production deployments of lightwave fabrics used for both datacenter networking and machine-learning (ML) applications. Using optical circuit switches (OCSes) and optical transceivers developed in-house, we employ hardware and software codesign to integrate the fabrics into our network and computing infrastructure. Key to our design is a high degree of multiplexing enabled by new kinds of wavelength-division-multiplexing (WDM) and optical circulators that support high-bandwidth bidirectional traffic on a single strand of optical fiber. The development of the requisite OCS and optical transceiver technologies leads to a synchronous lightwave fabric that is reconfigurable, low latency, rate agnostic, and highly available. These fabrics have provided substantial benefits for long-lived traffic patterns in our datacenter networks and predictable traffic patterns in tightly-coupled machine learning clusters. We report results for a large-scale ML superpod with 4096 tensor processing unit (TPU) V4 chips that has more than one ExaFLOP of computing power. For this use case, the deployment of a lightwave fabric provides up to 3× better system availability and model-dependent performance improvements of up to 3.3× compared to a static fabric, despite constituting less than 6% of the total system cost.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a0583de6-f712-4769-bd3d-d4e81a996008Cited by top-tier papers10
- Alibaba HPN: A Data Center Network for Large Language Model TrainingKun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao et al.SIGCOMM 2024 · 173 citations
- Astral: A Datacenter Infrastructure for Large Language Model Training at ScaleQingkai Meng, Hao Zheng, Zhenhui Zhang, ChonLam Lao et al.SIGCOMM 2025 · 16 citations
- Shale: A Practical, Scalable Oblivious Reconfigurable NetworkDaniel Amir, Nitika Saran, Tegan Wilson, Robert Kleinberg et al.SIGCOMM 2024 · 16 citations
- MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts TrainingXudong Liao, Yijun Sun, Han Tian, Xinchen Wan et al.SIGCOMM 2025 · 14 citations
- InfiniteHBD: Building Datacenter-Scale High-Bandwidth Domain for LLM with Optical Circuit Switching TransceiversChenchen Shou, Guyue Liu, Hao Nie, Huaiyu Meng et al.SIGCOMM 2025 · 9 citations
Related papers
- Reconfigurable Torus Fabrics for Multi-tenant MLAbhishek Vijaya Kumar, Eric Ding, Arjun Devraj, Darius Bunandar et al.ASPLOS 2026 · 1 citation
- Resiliency at Scale: Managing Google's TPUv4 Machine Learning SupercomputerYazhou Zu, Alireza Ghaffarkhah, Hoang-Vu Dang, Brian Towles et al.NSDI 2024 · 46 citations
- Opus: Photonic Rail-Optimized Fabric in ML DatacentersEric Ding, Barry Lyu, Bhaskar Kataria, Rachee SinghSIGCOMM 2026
- LightML: A Photonic Accelerator for Efficient General Purpose Machine LearningLiang Liu, Sadra Rahimi Kari, Xin Xin, Nathan Youngblood et al.ISCA 2025 · 3 citations
- A software-defined tensor streaming multiprocessor for large-scale machine learningDennis Abts, Garrin Kimmell, Andrew C. Ling, John Kim et al.ISCA 2022 · 46 citations
