Enabling AI Network Cross-Layer Design and Operations with Arcadia: A Simulation Platform at Scale
Zhaodong Wang, Satyajeet Singh Ahuja, Xu Zhang, Max Noormohammadpour, Gregory R. Steinbrecher, Thomas Fuller, Xin Liu, Kevin Quirk, Mikel Jimenez Fernandez, Abhinav Triguna, Yan Cai, Steve Politis
Abstract
The rapid evolution of Artificial Intelligence (AI) technology is fueling significant investments by hyperscalers, making AI networks crucial for large-scale training. Understanding the design impacts on AI training requires systematic, cross-layer evaluation. Production experience highlights the need for a robust simulation platform to guide network design and operations. This paper defines the platform requirements, addresses complex design challenges, and shares our experience building Arcadia, a scalable, high-fidelity simulation platform for AI Networks. It operates at the cluster level, focusing on overall cluster performance rather than individual job performance. By using our fast-forwarding, lock-free, and synchronizationcost reduction mechanisms, Arcadia achieves scalability and speed, allowing us to faithfully simulate real-world-scale training clusters and plays a key role in guiding Meta's AI network evolution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2244197b-7c36-4469-be16-9c3a6fddae56Builds on11
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi et al.OSDI 2020 · 390 citations
- RDMA over Ethernet for Distributed Training at Meta ScaleAdithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu et al.SIGCOMM 2024 · 171 citations
- Habitat: A Runtime-Based Computational Performance Predictor for Deep Neural Network TrainingGeoffrey X. Yu, Yubo Gao, Pavel Golikov, Gennady PekhimenkoUSENIX ATC 2021 · 108 citations
- Daydream: Accurately Estimating the Efficacy of Optimizations for DNN TrainingHongyu Zhu, Amar Phanishayee, Gennady PekhimenkoUSENIX ATC 2020 · 74 citations
- MimicNet: fast performance estimates for data center networks with machine learningQizhen Zhang, Kelvin K. W. Ng, Charles W. Kazer, Shen Yan et al.SIGCOMM 2021 · 63 citations
Related papers
- Nüwa: A Generative Control Plane for AI Network SimulationWenkai Li, Ran Shu, Peng Zhang, Yiren Zhao et al.SIGCOMM 2026
- Matryoshka: Realizing Hyperscale Data Center Network Design for the AI EraYan Cai, Jialong Li, Kutalmis Akpinar, Tianxiang Li et al.NSDI 2026 · 3 citations
- Accelerating Design Space Exploration for LLM Training Systems with Multi-experiment Parallel SimulationFei Gui, Kaihui Gao, Li Chen, Dan Li et al.NSDI 2025 · 27 citations
- SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and PrecisionXizheng Wang, Qingxu Li, Yichi Xu, Gang Lu et al.NSDI 2025 · 82 citations
- HeteroSim: Towards High-Fidelity Heterogeneous LLM Training Simulation on GPUsXiaofei Yue, Fangming Zhao, Fulun Ye, Jiongchi Yu et al.WWW 2026
