To PRI or Not To PRI, That's the question
Yun Wang, Liang Chen, Jie Ji, Xianting Tian, Ben Luo, Zhixiang Wei, Zhibai Huang, Kailiang Xu, Kaihuan Peng, Kaijie Guo, Ning Luo, Guangjian Wang
Abstract
SR-IOV and I/O device passthrough enable network and storage devices to be shared among multiple tenants with high density using virtual functions (VFs), achieving nearnative performance. However, passthrough does not support page faults, requiring the hypervisor to statically pin the VMallocated memory. This approach is unacceptable for cloud service providers (CSPs) that rely on oversubscription to enhance memory utilization and reduce costs. The Page Request Interface (PRI) was designed to support device-side I/O page faults (IOPFs) through collaboration among devices, Input-Output Memory Management Units (IOMMU), and the OS. But PRI has not seen broad adoption in devices like NICs and storage.
We propose VIO, a novel dynamic I/O device passthrough approach that achieves near-native performance and is hardware-independent. By leveraging a shadow available queue, VIO can dynamically and transparently switch devices between VIO and passthrough modes based on I/O operations per second (IOPS) pressure, balancing resource utilization and performance. Each DMA request is probed via IOPAsnooping in the virtio data plane to eliminate IOPFs, while device interrupts are directly passed through to the VM guest, enabling performance close to passthrough. VIO is extensively tested and deployed by a leading global CSP across 300K VMs, supporting both legacy and new instances while reclaiming up to the equivalent of 30K VM memory daily without compromising user Service Level Objectives (SLOs). As the scale grows, the benefits continue to increase.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Clio: a hardware-software co-designed disaggregated memory systemZhiyuan Guo, Yizhou Shan, Xuhao Luo, Yutong Huang et al.ASPLOS 2022 · 110 citations
- Exploring the Design Space of Page Management for Multi-Tiered Memory SystemsJonghyeon Kim, Wonkyo Choe, Jeongseob AhnUSENIX ATC 2021 · 108 citations
- TMO: transparent memory offloading in datacentersJohannes Weiner, Niket Agarwal, Dan Schatzberg, Leon Yang et al.ASPLOS 2022 · 103 citations
- Batch-Aware Unified Memory Management in GPUs for Irregular WorkloadsHyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi et al.ASPLOS 2020 · 89 citations
- Towards an Adaptable Systems Architecture for Memory Tiering at Warehouse-ScalePadmapriya Duraisamy, Wei Xu, Scott Hare, Ravi Rajwar et al.ASPLOS 2023 · 74 citations
Related papers
- VPRI: Efficient I/O Page Fault Handling via Software-Hardware Co-Design for IaaS CloudsKaijie Guo, Dingji Li, Ben Luo, Yibin Shen et al.SOSP 2024 · 1 citation
- HD-IOV: SW-HW Co-designed I/O Virtualization with Scalability and Flexibility for Hyper-Density CloudZongpu Zhang, Jiangtao Chen, Banghao Ying, Yahui Cao et al.EuroSys 2024 · 7 citations
- coIOMMU: A Virtual IOMMU with Cooperative DMA Buffer Tracking for Efficient Memory Management in Direct I/OKun Tian, Yu Zhang, Luwei Kang, Yan Zhao et al.USENIX ATC 2020 · 18 citations
- Efficient Memory Overcommitment for I/O Passthrough Enabled VMs via Fine-grained Page Meta-data ManagementYaohui Wang, Ben Luo, Yibin ShenUSENIX ATC 2023 · 18 citations
- NVMePass: A Lightweight, High-performance and Scalable NVMe Virtualization Architecture with I/O Queues PassthroughYiquan Chen, Zhen Jin, Yijing Wang, Yi Chen et al.HPCA 2025 · 1 citation
