Towards Exploiting CPU Elasticity via Efficient Thread Oversubscription
Hang Huang, Jia Rao, Song Wu, Hai Jin, Hong Jiang, Hao Che, Xiaofeng Wu
Abstract
Elasticity is an essential feature of cloud computing, which allows users to dynamically add or remove resources in response to workload changes. However, building applications that truly exploit elasticity is non-trivial. Traditional applications need to be modified to efficiently utilize variable resources. This paper explores thread oversubscription, i.e., provisioning more threads than the available cores, to exploit CPU elasticity in the cloud. While maintaining sufficient concurrency allows applications to utilize additional CPUs when more are made available, it is widely believed that thread oversubscription introduces prohibitive overheads due to excessive context switches, loss of locality, and contention on shared resources.
In this paper, we conduct a comprehensive study of the overhead of thread oversubscription. We find that 1) the direct cost of context switching (i.e., 1-2 𝜇𝑠 on modern processors) does not cause noticeable performance slow down to most applications; 2) oversubscription can be both constructive and destructive to the performance of CPU caches and TLB. We identify two previously under-studied issues that are responsible for drastic slowdowns in many applications under oversubscription. First, the existing thread sleep and wakeup process in the OS kernel is inefficient in handling oversubscribed threads. Second, pervasive busy-waiting operations in program code can waste CPU and starve critical threads. To this end, we devise two OS mechanisms, virtual blocking and busy-waiting detection, to enable efficient thread oversubscription without requiring program code changes. Experimental results show that our approaches can achieve an efficiency close to that in under-subscribed scenarios while preserving the capability to expand to many more CPUs. The performance gain is up to 77% for blocking-and 19x for busy-waiting-based applications compared to the vanilla Linux.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef92948b-7a1e-4318-ab76-ab77d5ccce40Cited by top-tier papers1
Ask how each one uses itRelated papers
- Computation Is Fast, Use T hreadlet !: Efficient Threading for μs-Scale Computing via OS/Hardware Co-DesignYiming Yao, Xiaohe Qin, Yi Fan, Yuanlong Li et al.SOSP 2026
- vSMT-IO: Improving I/O Performance and Efficiency on SMT Processors in Virtualized CloudsWeiwei Jia, Jianchen Shan, Tsz On Li, Xiaowei Shang et al.USENIX ATC 2020 · 16 citations
- UFO: The Ultimate QoS-Aware Core Management for Virtualized and Oversubscribed Public CloudsYajuan Peng, Shuang Chen, Yi Zhao, Zhibin YuNSDI 2024 · 7 citations
- A Theoretical Approach to Determine the Optimal Size of a Thread Pool for Real-Time SystemsDaniel CasiniRTSS 2022 · 5 citations
- Prediction-Based Power Oversubscription in Cloud PlatformsAlok Gautam Kumbhare, Reza Azimi, Ioannis Manousakis, Anand Bonde et al.USENIX ATC 2021 · 90 citations
