SMTcheck: Accurate SMT Interference Prediction to Improve Scheduling Efficiency in Datacenters
Sanghyun Kim, Jinhyeok Oh, Taehun Kim, Gyutae Kim, Youngsok Kim, Jaehyun Hwang, Joonsung Kim
Abstract
Simultaneous multithreading (SMT) is widely used in modern x86 processors to improve core utilization by sharing hardware resources between co-located threads. However, such resource sharing often leads to severe performance interference, making efficient workload co-scheduling difficult, especially given the complexity and diversity of modern x86 CPUs. Our analysis reveals that SMT-aware workload scheduling can significantly improve system throughput and reduce tail latency for datacenter workloads, but identifying optimal thread combinations is challenging due to the lack of visibility into platform-specific resource sharing behaviors. In this paper, we present SMTcheck, a lightweight, accurate, and platformindependent methodology for predicting SMT interference for diverse x86 processors. SMTcheck uses carefully designed code snippets (Diags) to extract hidden microarchitectural features of performance-critical shared resources. With these extracted features, SMTcheck builds per-resource microbenchmarks (Injectors) to apply pinpoint pressure to specific target resources in order to capture workload-specific contention characteristics. SMTcheck then constructs a hardware-aware contention model to predict performance interference between arbitrary workload pairs without requiring exhaustive profiling. We evaluate SMTcheck on six x86 desktop processors and five x86 server processors from Intel and AMD across different generations and show that it achieves high prediction accuracy by up to 95.5 % (94.6 % on average). We further demonstrate its effectiveness by implementing a contention-aware scheduler in the Linux kernel. Compared to the default Linux scheduler, our contention-aware scheduler significantly reduces tail latency for latency-critical workloads (e.g., database, key-value store) by up to 36.09 %, and improves the overall system throughput by up to. Finally, using real-world cluster traces from Alibaba and Google, we demonstrate that SMTcheck incurs negligible profiling overheads (), making it practical for deployment in productionscale datacenter environments.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Ghost Threading: Helper-Thread Prefetching for Real SystemsYuxin Guo, Akshay Bhosale, Utpal Bora, Alexandra W. Chadwick et al.MICRO 2025 · 2 citations
- Holmes: SMT Interference Diagnosis and CPU Scheduling for Job Co-locationAidi Pi, Xiaobo Zhou, Chengzhong XuHPDC 2022 · 11 citations
- Harvesting Memory-bound CPU Stall Cycles in Software with MSHZhihong Luo, Sam Son, Sylvia Ratnasamy, Scott ShenkerOSDI 2024 · 5 citations
- vSMT-IO: Improving I/O Performance and Efficiency on SMT Processors in Virtualized CloudsWeiwei Jia, Jianchen Shan, Tsz On Li, Xiaowei Shang et al.USENIX ATC 2020 · 16 citations
- LibPreemptible: Enabling Fast, Adaptive, and Hardware-Assisted User-Space SchedulingYueying Li, Nikita Lazarev, David Koufaty, Tenny Yin et al.HPCA 2024 · 14 citations
