Alibaba Stellar: A New Generation RDMA Network for Cloud AI
Jie Lu, Jiaqi Gao, Fei Feng, Zhiqiang He, Menglei Zheng, Kun Liu, Jun He, Binbin Liao, Suwei Xu, Ke Sun, Yongjia Mo, Qinghua Peng
Abstract
The rapid adoption of Large Language Models (LLMs) in cloud environments has intensified the demand for high-performance AI training and inference, where Remote Direct Memory Access (RDMA) plays a critical role. However, existing RDMA virtualization solutions, such as Single-Root Input/Output Virtualization (SR-IOV), face significant limitations in scalability, performance, and stability. These issues include lengthy container initialization times, hardware resource constraints, and inefficient traffic steering. To address these challenges, we propose Stellar, a new generation RDMA network for cloud AI. Stellar introduces three key innovations: Para-Virtualized Direct Memory Access (PVDMA) for on-demand memory pinning, extended Memory Translation Table (eMTT) for optimized GPU Direct RDMA (GDR) performance, and RDMA Packet Spray for efficient multi-path utilization. Deployed in our large-scale AI clusters, Stellar spins up virtual devices in seconds, reduces container initialization time by 15 times, and improves LLM training speed by up to 14%. Our evaluations demonstrate that Stellar significantly outperforms existing solutions, offering a scalable, stable, and high-performance RDMA network for cloud AI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Efficient and Flexible Datapaths for Fine-Grained Rack-Scale Interconnects with Elastic QPChenxingyu Zhao, Yibo Wu, Hongtao Zhang, Jaehong Min et al.SIGCOMM 2026 · 1 citation
- STORM: Enabling Traffic Scheduling for RDMAJichun Wu, Ran Shu, Gianni Antichi, Yongqiang Xiong et al.SIGCOMM 2026
- Bifrost: Alibaba's Next-Generation VPC Network with High-Performance Multipath Reliable TransportZihao Fan, Xing Li, Ye Yang, Bo Jiang et al.NSDI 2026
Builds on6
- Alibaba HPN: A Data Center Network for Large Language Model TrainingKun Qian, Yongqing Xi, Jiamin Cao, Jiaqi Gao et al.SIGCOMM 2024 · 173 citations
- RDMA over Ethernet for Distributed Training at Meta ScaleAdithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu et al.SIGCOMM 2024 · 171 citations
- Network Load Balancing with In-network Reordering Support for RDMACha Hwan Song, Xin Zhe Khooi, Raj Joshi, Inho Choi et al.SIGCOMM 2023 · 110 citations
- RunD: A Lightweight Secure Container Runtime for High-density Deployment and High-concurrency Startup in Serverless ComputingZijun Li, Jiagan Cheng, Quan Chen, Eryu Guan et al.USENIX ATC 2022 · 106 citations
- MasQ: RDMA for Virtual Private CloudZhiqiang He, Dongyang Wang, Binzhang Fu, Kun Tan et al.SIGCOMM 2020 · 38 citations
Related papers
- RDNet: An RDMA-aware Container Network Interface for Cloud EnvironmentsMyoungsung You, Minjae Seo, Seungwon Shin, Jaehyun NamINFOCOM 2026
- SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA CommunicationMikhail Khalilov, Siyuan Shen, Marcin Chrapek, Tiancheng Chen et al.SC 2025 · 6 citations
- Vela: A Virtualized LLM Training System with GPU Direct RoCEApoorve Mohan, Robert Walkup, Bengi Karacali, Ming-Hung Chen et al.ASPLOS 2025 · 3 citations
- SimAI: Unifying Architecture Design and Performance Tuning for Large-Scale Large Language Model Training with Scalability and PrecisionXizheng Wang, Qingxu Li, Yichi Xu, Gang Lu et al.NSDI 2025 · 82 citations
- Tlaloc: A Generic Multipath Load Balancing for RoCEHuimin Luo, Jiao Zhang, Yongchen Pan, Tian Pan et al.INFOCOM 2026
