HeteroSim: Towards High-Fidelity Heterogeneous LLM Training Simulation on GPUs
Xiaofei Yue, Fangming Zhao, Fulun Ye, Jiongchi Yu, Zhaoxuan Li, Tingting Li, Ziming Zhao, Jianwei Yin
摘要
Modern Large Language Model (LLM) training clusters increasingly mix heterogeneous GPUs, diverse intra-node fabrics, and inter-node interconnects, combined with varied parallelism strategies. Exploring this massive design space, further amplified by heterogeneity, through real deployments is prohibitively slow and costly. Existing simulators, which are primarily designed and tuned for homogeneous clusters, either trade fidelity for speed or require heavyweight workflows with non-negligible overhead. We propose HeteroSim, a high-fidelity simulation framework for heterogeneous LLM training systems. It introduces: (i) a LLM training workload compiler that captures realistic training graphs, microbatching schedules, and compute-communication overlap; (ii) a heterogeneity-aware computation planner using roofline-style scaling across GPU generations; (iii) a collective communication planner that reproduces NCCL-like behaviors with per-link models, message channelization, and configurable routing. Across a wide range of heterogeneity levels, experimental results show that HeteroSim achieves near-real simulation accuracy while keeping low overhead.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and PrecisionXizheng Wang, Qingxu Li, Yichi Xu, Gang Lu 等NSDI 2025 · 被引用 82 次
- HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for Heterogeneous GPU ClustersChenyang Hei, Jiayi Li, Jiamin Cao, Chengxi Gao 等NSDI 2026 · 被引用 4 次
- A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with CrystalLLMShaoke Xi, ChonLam Lao, Boyi Jia, Jiaqi Gao 等SOSP 2026
- Scalable Synthesis of Distributed Llm Workloads Through Symbolic Tensor GraphsChanghai Man, Joongun Park, Hanjiang Wu, Huan Xu 等ISCA 2026 · 被引用 2 次
- Accelerating Design Space Exploration for LLM Training Systems with Multi-experiment Parallel SimulationFei Gui, Kaihui Gao, Li Chen, Dan Li 等NSDI 2025 · 被引用 27 次
