SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and Precision
Xizheng Wang, Qingxu Li, Yichi Xu, Gang Lu, Dan Li, Li Chen, Heyang Zhou, Linkang Zheng, Sen Zhang, Yikai Zhu, Yang Liu, Pengcheng Zhang
Abstract
The large number of GPUs required for a single LLM training significantly hinders the validation of new designs, tunings, and optimizations, calling for the occurrence of efficient simulators. Existing simulators, however, only target a specific granularity of the entire training, intrinsically leading to imprecision. This paper presents SimAI, a unified simulator aiming at precisely and efficiently simulating the LLM training procedure at scale. Through selective and high-fidelity integration of the training frameworks, the kernel computation, and the collective communication library into the simulating procedure, SimAI achieves high precision in simulations. SimAI further conducts multi-thread acceleration and implements lock-free global context-sharing to accelerate the execution speed. The effectiveness of SimAI is validated by its performance results, which show an average of 98.1% alignment to real-world results under various test scenarios and affirm its robustness and adaptability from small-scale labs to large-scale industrial environments. SimAI delivers meaningful guidelines for new host designs and parameter settings, directly benefiting in-production LLM training. We also share experiences and lessons learned during the evolution of SimAI. SimAI is open sourced at https://github.com/aliyun/SimAI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb675f48-a300-437d-8d9f-6ff36ae9d8f9Cited by top-tier papers11
- Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance EstimationJianxing Qin, Jingrong Chen, Xinhao Kong, Yongji Wu et al.NSDI 2026 · 5 citations
- Compass: SLO-aware Query Planner for Compound AI Serving at ScaleBanruo Liu, Wei-Yu Lin, Minghao Fang, Yihan Jiang et al.VLDB 2026 · 5 citations
- Supercharging Packet-level Network Simulation of Large Model Training via Memoization and Fast-ForwardingFei Long, Kaihui Gao, Li Chen, Dan Li et al.NSDI 2026 · 4 citations
- Scalable Synthesis of Distributed Llm Workloads Through Symbolic Tensor GraphsChanghai Man, Joongun Park, Hanjiang Wu, Huan Xu et al.ISCA 2026 · 2 citations
- PipeWeave: Synergizing Analytical and Learning Models for Unified GPU Performance PredictionKaixuan Zhang, Yunfan Cui, Shuhao Zhang, Chutong Ding et al.ISCA 2026 · 2 citations
Builds on11
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- Accel-Sim: An Extensible Simulation Framework for Validated GPU ModelingMahmoud Khairy, Zhesheng Shen, Tor M. Aamodt, Timothy G. RogersISCA 2020 · 366 citations
- Alibaba HPN: A Data Center Network for Large Language Model TrainingKun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao et al.SIGCOMM 2024 · 173 citations
- Collie: Finding Performance Anomalies in RDMA SubsystemsXinhao Kong, Yibo Zhu, Huaping Zhou, Zhuo Jiang et al.NSDI 2022 · 86 citations
- From luna to solar: the evolutions of the compute-to-storage networks in Alibaba cloudRui Miao, Lingjun Zhu, Shu Ma, Kun Qian et al.SIGCOMM 2022 · 79 citations
Related papers
- HeteroSim: Towards High-Fidelity Heterogeneous LLM Training Simulation on GPUsXiaofei Yue, Fangming Zhao, Fulun Ye, Jiongchi Yu et al.WWW 2026
- vTrain: A Simulation Framework for Evaluating Cost-Effective and Compute-Optimal Large Language Model TrainingJehyeon Bang, Yujeong Choi, Myeongwoo Kim, Yongdeok Kim et al.MICRO 2024 · 16 citations
- Accelerating Design Space Exploration for LLM Training Systems with Multi-experiment Parallel SimulationFei Gui, Kaihui Gao, Li Chen, Dan Li et al.NSDI 2025 · 27 citations
- A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with CrystalLLMShaoke Xi, ChonLam Lao, Boyi Jia, Jiaqi Gao et al.SOSP 2026
- ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM TrainingAdel Nabli, Louis Fournier, Pierre Erbacher, Louis Serrano et al.NeurIPS 2025 · 5 citations
