UFO: The Ultimate QoS-Aware Core Management for Virtualized and Oversubscribed Public Clouds
Yajuan Peng, Shuang Chen, Yi Zhao, Zhibin Yu
Abstract
Public clouds typically adopt (1) multi-tenancy to increase server utilization; (2) virtualization to provide isolation between different tenants; (3) oversubscription of resources to further increase resource efficiency. However, prior work all focuses on optimizing one or two elements, and fails to considerately bring QoS-aware multi-tenancy, virtualization and resource oversubscription together.
We find three challenges when the three elements coexist. First, the double scheduling symptoms are 10× worse with latency-critical (LC) workloads which are comprised of numerous sub-millisecond tasks and are significantly different from conventional batch applications. Second, inner-VM resource contention also exists between threads of the same VM when running LC applications, calling for inner-VM core isolation. Third, no application-level performance metrics can be obtained by the host to guide resource management in realistic public clouds.
To address these challenges, we propose a QoS-aware core manager dubbed UFO to specifically support co-location of multiple LC workloads in virtualized and oversubscribed public cloud environments. UFO solves the three abovementioned challenges, by (1) coordinating the guest and host CPU cores (vCPU-pCPU coordination), and (2) doing finegrained inner-VM resource isolation, to push core management in realistic public clouds to the extreme. Compared with the state-of-the-art core manager, it saves up to 50% (average of 22%) of physical cores under the same co-location scenario.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7dd81f8a-f0c0-455e-971e-4cb293ab404dCited by top-tier papers5
- Improving GPU Sharing Performance through Adaptive Bubbleless Spatial-Temporal SharingShulai Zhang, Quan Chen, Weihao Cui, Han Zhao et al.EuroSys 2025 · 19 citations
- Coach: Exploiting Temporal Patterns for All-Resource Oversubscription in Cloud PlatformsBenjamin Reidys, Pantea Zardoshti, Íñigo Goiri, Celine Irvene et al.ASPLOS 2025 · 9 citations
- Enabling SLO-Aware 5G Multi-Access Edge Computing with SMECXiao Zhang, Daehyeok KimNSDI 2026 · 4 citations
- Optimizing Task Scheduling in Cloud VMs with Accurate vCPU AbstractionEdward Guo, Weiwei Jia, Xiaoning Ding, Jianchen ShanEuroSys 2025 · 3 citations
- Svalinn: Overload Control in Large-Scale Servers with Multiple Resource BottlenecksBhaskar Subhash Pardeshi, Peidi Song, Ahmed SaeedOSDI 2026
Builds on5
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale ComputersTirthak Patel, Devesh TiwariHPCA 2020 · 153 citations
- Twig: Multi-Agent Task Management for Colocated Latency-Critical Cloud ServicesRajiv Nishtala, Vinicius Petrucci, Paul M. Carpenter, Magnus SjälanderHPCA 2020 · 76 citations
- ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the CloudShuang Chen, Angela Jin, Christina Delimitrou, José F. MartínezHPCA 2022 · 34 citations
- Ah-Q: Quantifying and Handling the Interference within a Datacenter from a System PerspectiveYuhang Liu, Xin Deng, Jiapeng Zhou, Mingyu Chen et al.HPCA 2023 · 15 citations
Related papers
- OLPart: Online Learning based Resource Partitioning for Colocating Multiple Latency-Critical Jobs on Commodity ComputersRuobing Chen, Haosen Shi, Yusen Li, Xiaoguang Liu et al.EuroSys 2023 · 27 citations
- Rhythm: component-distinguishable workload deployment in datacentersLaiping Zhao, Yanan Yang, Kaixuan Zhang, Xiaobo Zhou et al.EuroSys 2020 · 49 citations
- EcoCore: Dynamic Core Management for Improving Energy Efficiency in Latency-Critical ApplicationsGyeongseo Park, Minho Kim, Ki-Dong Kang, Yunhyeong Jeon et al.MICRO 2025 · 1 citation
- You Can Always Get What You Want: CPU Virtualization Made Fast and FreeYun Wang, Xingguo Jia, Ben Luo, Kenan Liu et al.SOSP 2026
- SmartHarvest: harvesting idle CPUs safely and efficiently in the cloudYawen Wang, Kapil Arya, Marios Kogias, Manohar Vanga et al.EuroSys 2021 · 53 citations
