Astral: A Datacenter Infrastructure for Large Language Model Training at Scale
Qingkai Meng, Hao Zheng, Zhenhui Zhang, ChonLam Lao, Chengyuan Huang, Baojia Li, Ziyuan Zhu, Hao Lu, Weizhen Dang, Zitong Lin, Weifeng Zhang, Lingfeng Liu
Abstract
The flourishing of Large Language Models (LLMs) calls for increasingly ultra-scale training. In this paper, we share our experience in designing, deploying, and operating our novel Astral datacenter infrastructure, along with operational lessons and evolutionary insights gained from its production use. Astral has three important innovations: (i) a same-rail interconnection network architecture on tier-2, which enables the scaling of LLM training. To physically deploy this high-density infrastructure, we introduce a distributed high-voltage direct current power system and a new air-liquid integrated cooling system. (ii) a full-stack monitoring system featuring cross-host and hierarchical logging correlation, which diagnoses failures at scale and precisely localizes root causes. (iii) an operator-granular forecasting component Seer that efficiently generates operator execution timelines with acceptable accuracy, aiding in fault diagnosis, model tuning, and network architecture upgrading. Astral infrastructure has been gradually deployed over 18 months, supporting LLM training and inference for multiple customers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1f404e2f-5c61-49a6-9044-64ca4f8c0c5cCited by top-tier papers5
- SwiftEP: Accelerating MoE Inference with Buffer Fusion and TMA OffloadingXingyi Li, Yadong Liu, Xiaojie Huang, Yiran Zhang et al.NSDI 2026 · 2 citations
- EPIC: Abstraction and Polymorphism of In-Network Collectives on EthernetYitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou et al.SIGCOMM 2026 · 1 citation
- Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and InferenceMostafa Elhoushi, Alexander Pretko, Nolan Dey, Bin Zhang et al.ICML 2026
- TrainMover: An Interruption-Resilient Runtime for ML TrainingChonLam Lao, Jiaqi Gao, Jiamin Cao, Zhipeng Zhang et al.OSDI 2026
- Fine-grained and Non-intrusive LLM Training Monitoring via Microsecond-level Traffic MeasurementYibo Xiao, Hao Zheng, Haifeng Sun, Qingkai Meng et al.ASPLOS 2026
Builds on21
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
Related papers
- Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMsHouyi Li, Wenzhen Zheng, Qiufeng Wang, Zhenyu Ding et al.NeurIPS 2025 · 4 citations
- Robust LLM Training Infrastructure at ByteDanceBorui Wan, Gaohong Liu, Zuquan Song, Jun Wang et al.SOSP 2025 · 1 citation
- MegaScale: Scaling Large Language Model Training to More Than 10, 000 GPUsZiheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang et al.NSDI 2024 · 415 citations
- Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUsGuoliang He, Youhe Jiang, Wencong Xiao, Kaihua Jiang et al.NeurIPS 2025 · 10 citations
- Di-PS: System-Algorithm Co-Design for Asynchronous and Heterogeneous Cross-cluster LLM Training at ScaleShengwei Li, Qiaoling Chen, Zhiquan Lai, Penglong Jiao et al.NSDI 2026
