CoPilotIO: CPU as a Co-Pilot for GPU I/O to Free GPU Compute
Guanyi Chen, Qi Chen, Shu Yin, Jian Zhang
Abstract
Limited GPU memory increasingly forces modern AI and data analytics workloads to access terabyte-scale datasets and model states from storage, making efficient GPU I/O critical. Existing GPU I/O engines are either CPU-centric or GPU-centric. CPU-centric approaches avoid consuming GPU resources but often fail to provide high-throughput, ondemand GPU access due to kernel overheads and limited parallelism. GPU-centric approaches enable fine-grained ondemand I/O but require intensive I/O polling that consumes valuable GPU resources and introduces intra-warp, inter-warp, and inter-SM I/O stalls. We present CoPilotIO 1 , a novel GPU I/O engine that delivers high-throughput, on-demand storage access without sacrificing GPU compute resources. CoPilo-tIO adopts an asynchronous GPU I/O architecture in which GPUs initiate I/O while CPU cores act as I/O co-pilots responsible for completion polling. To enable efficient coordination, CoPilotIO introduces a split SQ/CQ architecture, hardware barrier-based synchronization, a lock-free barrier-table, and adaptive CPU-GPU co-polling. Across microbenchmarks and real applications, including GoFS, LLM Mixture-of-Experts (MoE) inference, and Deep Learning Recommendation Models (DLRM), CoPilotIO reduces I/O-induced stalls by up to 55.5%, requires 50% fewer SMs to saturate the GPU PCIe bandwidth, accelerates GoFS by up to 17.4%, and improves application performance by up to 85%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64401ebc-8bbb-4fbf-8bd7-a7168a16434aBuilds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li et al.ICML 2023 · 683 citations
- EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal In GPUsSeungwon Min, Vikram Sharma Mailthody, Zaid Qureshi, Jinjun Xiong et al.VLDB 2021 · 66 citations
- GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System ArchitectureZaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado, Seungwon Min et al.ASPLOS 2023 · 48 citations
Related papers
- AGILE: Lightweight and Efficient Asynchronous GPU-SSD IntegrationZhuoping Yang, Jinming Zhuang, Xingzhen Chen, Alex K. Jones et al.SC 2025 · 3 citations
- Asynchrony and GPUs: Bridging this Dichotomy for I/O with AGIOJihoon Han, Anand Sivasubramaniam, Chia-Hao Chang, Vikram Sharma Mailthody et al.ASPLOS 2026 · 1 citation
- Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data AnalyticsYichao Yuan, Advait Iyer, Lin Ma, Nishil TalatiVLDB 2025 · 11 citations
- Managing Scalable Direct Storage Accesses for GPUs with GoFSShaobo Li, Yirui Eric Zhou, Yuqi Xue, Yuan Xu et al.SOSP 2025
- PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast StorageJigao Luo, Nils Boeschen, Muhammad El-Hindi, Carsten BinnigVLDB 2026
