Fast Cloud Storage for AI Jobs via Grouped I/O API with Transparent Read/Write Optimizations
Yingyi Hao, Ting Yao, Xingda Wei, Dingyan Zhang, Tianle Sun, Yiwen Zhang, Zhiyong Fu, Huatao Wu, Rong Chen
Abstract
The emergence of AI workloads has placed rigorous bandwidth requirements on cloud storage, which are challenging to meet due to inherent hardware restrictions in cost-efficient disaggregated storage architectures, as well as the non-triviality of implementing application-tailored optimizations.
This paper presents AITURBO, a cloud storage system for AI jobs with high bandwidth demands. AITURBO first utilizes the high-bandwidth compute fabric between accelerators to meet AI applications' bandwidth demands without incurring additional storage cost. AITURBO further introduces a simple yet powerful grouped I/O API that allows AITURBO to automatically derive optimized read and write plans at the storage layer. These plans enable optimizations that are comparable or better than application-level ones, because they capture common I/O patterns in AI workloads and have a holistic view from the storage layer's perspective. Under common AI workloads such as checkpoint reads and writes and KV-cache reads, AITURBO achieves comparable or better performance than state-of-the-art systems, with and without application-level optimizations, including systems such as Megatron, Gemini, and Mooncake, typically with minimal application-level code changes. AITURBO has been deployed in training jobs in HUAWEI's production cloud to support efficient training workloads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext caa4222b-d284-4bb1-b180-e1628d8c273aBuilds on17
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- Characterization of Large Language Model Development in the DatacenterQinghao Hu, Zhisheng Ye, Zerui Wang, Guoteng Wang et al.NSDI 2024 · 192 citations
- CheckFreq: Frequent, Fine-Grained DNN CheckpointingJayashree Mohan, Amar Phanishayee, Vijay ChidambaramFAST 2021 · 175 citations
- Facebook's Tectonic Filesystem: Efficiency from ExascaleSatadru Pan, Theano Stavrinos, Yunqiao Zhang, Atul Sikaria et al.FAST 2021 · 110 citations
Related papers
- PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast StorageJigao Luo, Nils Boeschen, Muhammad El-Hindi, Carsten BinnigVLDB 2026
- AIIO: Using Artificial Intelligence for Job-Level and Automatic I/O Performance Bottleneck DiagnosisBin Dong, Jean Luca Bez, Suren BynaHPDC 2023 · 5 citations
- Dynamic Resource Allocation for Deep Learning Clusters with Separated Compute and StorageMingxia Li, Zhenhua Han, Chi Zhang, Ruiting Zhou et al.INFOCOM 2023 · 3 citations
- SmartDS: Middle-Tier-centric SmartNIC Enabling Application-aware Message Split for Disaggregated Block StorageJie Zhang, Hongjing Huang, Lingjun Zhu, Shu Ma et al.ISCA 2023 · 15 citations
- ArrayMorph: Optimizing Hyperslab Queries on the Cloud for Machine Learning PipelinesRuochen Jiang, Spyros BlanasVLDB 2025
