ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
Siyuan Shen, Tommaso Bonato, Zhiyi Hu, Pasquale Jordan, Tiancheng Chen, Torsten Hoefler
摘要
Network simulators play a crucial role in evaluating the performance of large-scale systems. However, existing simulators rely heavily on synthetic microbenchmarks or narrowly focus on specific domains, limiting their ability to provide comprehensive performance insights. In this work, we introduce ATLAHS, a flexible, extensible, and open-source toolchain designed to trace real-world applications and accurately simulate their workloads. ATLAHS leverages the Group Operation Assembly Language (GOAL) format to model communication and computation patterns in AI, HPC, and distributed storage applications. It supports multiple network simulation backends and handles multi-job and multi-tenant scenarios. Through extensive validation, we demonstrate that ATLAHS achieves high accuracy in simulating realistic workloads (consistently less than 5% error), while significantly outperforming AstraSim, the current state-of-the-art AI systems simulator, in terms of both simulation runtime and trace size efficiency. We further illustrate ATLAHS’s utility via detailed case studies, highlighting the impact of congestion control algorithms on the performance of distributed storage systems, as well as the influence of job-placement strategies on application runtimes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel 等SIGCOMM 2020 · 被引用 333 次
- Exploring GPU-to-GPU Communication: Insights into Supercomputer InterconnectsDaniele De Sensi, Lorenzo Pichetti, Flavio Vella, Tiziano De Matteis 等SC 2024 · 被引用 23 次
- Evaluating Chiplet-based Large-Scale Interconnection Networks via Cycle-Accurate Packet-Parallel SimulationYinxiao Feng, Yuchen Wei, Dong Xiang, Kaisheng MaUSENIX ATC 2024 · 被引用 21 次
- HammingMesh: A Network Topology for Large-Scale Deep LearningTorsten Hoefler, Tommaso Bonato, Daniele De Sensi, Salvatore Di Girolamo 等SC 2022 · 被引用 20 次
- REPS: Recycled Entropy Packet Spraying for Adaptive Load Balancing and Failure MitigationTommaso Bonato, Abdul Kabbani, Ahmad Ghalayini, Michael Papamichael 等EuroSys 2026 · 被引用 12 次
相关 Paper
- Scalable Synthesis of Distributed Llm Workloads Through Symbolic Tensor GraphsChanghai Man, Joongun Park, Hanjiang Wu, Huan Xu 等ISCA 2026 · 被引用 2 次
- HeteroSim: Towards High-Fidelity Heterogeneous LLM Training Simulation on GPUsXiaofei Yue, Fangming Zhao, Fulun Ye, Jiongchi Yu 等WWW 2026
- Zettafly: A Network Topology with Flexible Non-blocking Regions for Large-scale AI and HPC SystemsDezun Dong, Ziyu Wang, Fei LeiISCA 2025 · 被引用 7 次
- LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear ProgrammingSiyuan Shen, Langwen Huang, Marcin Chrapek, Timo Schneider 等SC 2024 · 被引用 7 次
- Nüwa: A Generative Control Plane for AI Network SimulationWenkai Li, Ran Shu, Peng Zhang, Yiren Zhao 等SIGCOMM 2026
