Enabling AI Network Cross-Layer Design and Operations with Arcadia: A Simulation Platform at Scale
Zhaodong Wang, Satyajeet Singh Ahuja, Xu Zhang, Max Noormohammadpour, Gregory R. Steinbrecher, Thomas Fuller, Xin Liu, Kevin Quirk, Mikel Jimenez Fernandez, Abhinav Triguna, Yan Cai, Steve Politis
摘要
The rapid evolution of Artificial Intelligence (AI) technology is fueling significant investments by hyperscalers, making AI networks crucial for large-scale training. Understanding the design impacts on AI training requires systematic, cross-layer evaluation. Production experience highlights the need for a robust simulation platform to guide network design and operations. This paper defines the platform requirements, addresses complex design challenges, and shares our experience building Arcadia, a scalable, high-fidelity simulation platform for AI Networks. It operates at the cluster level, focusing on overall cluster performance rather than individual job performance. By using our fast-forwarding, lock-free, and synchronizationcost reduction mechanisms, Arcadia achieves scalability and speed, allowing us to faithfully simulate real-world-scale training clusters and plays a key role in guiding Meta's AI network evolution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU ClustersYimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi 等OSDI 2020 · 被引用 390 次
- RDMA over Ethernet for Distributed Training at Meta ScaleAdithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu 等SIGCOMM 2024 · 被引用 171 次
- Habitat: A Runtime-Based Computational Performance Predictor for Deep Neural Network TrainingGeoffrey X. Yu, Yubo Gao, Pavel Golikov, Gennady PekhimenkoUSENIX ATC 2021 · 被引用 108 次
- Daydream: Accurately Estimating the Efficacy of Optimizations for DNN TrainingHongyu Zhu, Amar Phanishayee, Gennady PekhimenkoUSENIX ATC 2020 · 被引用 74 次
- MimicNet: fast performance estimates for data center networks with machine learningQizhen Zhang, Kelvin K. W. Ng, Charles W. Kazer, Shen Yan 等SIGCOMM 2021 · 被引用 63 次
相关 Paper
- Nüwa: A Generative Control Plane for AI Network SimulationWenkai Li, Ran Shu, Peng Zhang, Yiren Zhao 等SIGCOMM 2026
- Matryoshka: Realizing Hyperscale Data Center Network Design for the AI EraYan Cai, Jialong Li, Kutalmis Akpinar, Tianxiang Li 等NSDI 2026 · 被引用 3 次
- Accelerating Design Space Exploration for LLM Training Systems with Multi-experiment Parallel SimulationFei Gui, Kaihui Gao, Li Chen, Dan Li 等NSDI 2025 · 被引用 27 次
- SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and PrecisionXizheng Wang, Qingxu Li, Yichi Xu, Gang Lu 等NSDI 2025 · 被引用 82 次
- HeteroSim: Towards High-Fidelity Heterogeneous LLM Training Simulation on GPUsXiaofei Yue, Fangming Zhao, Fulun Ye, Jiongchi Yu 等WWW 2026
