SC2025Top-tier venue
ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
Siyuan Shen, Tommaso Bonato, Zhiyi Hu, Pasquale Jordan, Tiancheng Chen, Torsten Hoefler
Abstract
Network simulators play a crucial role in evaluating the performance of large-scale systems. However, existing simulators rely heavily on synthetic microbenchmarks or narrowly focus on specific domains, limiting their ability to provide comprehensive performance insights. In this work, we introduce ATLAHS, a flexible, extensible, and open-source toolchain designed to trace real-world applications and accurately simulate their workloads. ATLAHS leverages the Group Operation Assembly Language (GOAL) format to model communication and computation patterns in AI, HPC, and distributed storage applications. It supports multiple network simulation backends and handles multi-job and multi-tenant scenarios. Through extensive validation, we demonstrate that ATLAHS achieves high accuracy in simulating realistic workloads (consistently less than 5% error), while significantly outperforming AstraSim, the current state-of-the-art AI systems simulator, in terms of both simulation runtime and trace size efficiency. We further illustrate ATLAHS’s utility via detailed case studies, highlighting the impact of congestion control algorithms on the performance of distributed storage systems, as well as the influence of job-placement strategies on application runtimes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 55bb2e8d-c55d-4c73-b63e-1c18ed88ccffCited by top-tier papers1
Ask how each one uses itBuilds on8
- Swift: Delay is Simple and Effective for Congestion Control in the DatacenterGautam Kumar, Nandita Dukkipati, Keon Jang, Hassan M. G. Wassel et al.SIGCOMM 2020 · 333 citations
- Exploring GPU-to-GPU Communication: Insights into Supercomputer InterconnectsDaniele De Sensi, Lorenzo Pichetti, Flavio Vella, Tiziano De Matteis et al.SC 2024 · 23 citations
- Evaluating Chiplet-based Large-Scale Interconnection Networks via Cycle-Accurate Packet-Parallel SimulationYinxiao Feng, Yuchen Wei, Dong Xiang, Kaisheng MaUSENIX ATC 2024 · 21 citations
- HammingMesh: A Network Topology for Large-Scale Deep LearningTorsten Hoefler, Tommaso Bonato, Daniele De Sensi, Salvatore Di Girolamo et al.SC 2022 · 20 citations
- REPS: Recycled Entropy Packet Spraying for Adaptive Load Balancing and Failure MitigationTommaso Bonato, Abdul Kabbani, Ahmad Ghalayini, Michael Papamichael et al.EuroSys 2026 · 12 citations
Related papers
- Scalable Synthesis of Distributed Llm Workloads Through Symbolic Tensor GraphsChanghai Man, Joongun Park, Hanjiang Wu, Huan Xu et al.ISCA 2026 · 2 citations
- HeteroSim: Towards High-Fidelity Heterogeneous LLM Training Simulation on GPUsXiaofei Yue, Fangming Zhao, Fulun Ye, Jiongchi Yu et al.WWW 2026
- Zettafly: A Network Topology with Flexible Non-blocking Regions for Large-scale AI and HPC SystemsDezun Dong, Ziyu Wang, Fei LeiISCA 2025 · 7 citations
- LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear ProgrammingSiyuan Shen, Langwen Huang, Marcin Chrapek, Timo Schneider et al.SC 2024 · 7 citations
- Nüwa: A Generative Control Plane for AI Network SimulationWenkai Li, Ran Shu, Peng Zhang, Yiren Zhao et al.SIGCOMM 2026
