Lune

HPCA2026Top-tier venue

SMTcheck: Accurate SMT Interference Prediction to Improve Scheduling Efficiency in Datacenters

Sanghyun Kim, Jinhyeok Oh, Taehun Kim, Gyutae Kim, Youngsok Kim, Jaehyun Hwang, Joonsung Kim

2026Year

Abstract

Simultaneous multithreading (SMT) is widely used in modern x86 processors to improve core utilization by sharing hardware resources between co-located threads. However, such resource sharing often leads to severe performance interference, making efficient workload co-scheduling difficult, especially given the complexity and diversity of modern x86 CPUs. Our analysis reveals that SMT-aware workload scheduling can significantly improve system throughput and reduce tail latency for datacenter workloads, but identifying optimal thread combinations is challenging due to the lack of visibility into platform-specific resource sharing behaviors. In this paper, we present SMTcheck, a lightweight, accurate, and platformindependent methodology for predicting SMT interference for diverse x86 processors. SMTcheck uses carefully designed code snippets (Diags) to extract hidden microarchitectural features of performance-critical shared resources. With these extracted features, SMTcheck builds per-resource microbenchmarks (Injectors) to apply pinpoint pressure to specific target resources in order to capture workload-specific contention characteristics. SMTcheck then constructs a hardware-aware contention model to predict performance interference between arbitrary workload pairs without requiring exhaustive profiling. We evaluate SMTcheck on six x86 desktop processors and five x86 server processors from Intel and AMD across different generations and show that it achieves high prediction accuracy by up to 95.5 % (94.6 % on average). We further demonstrate its effectiveness by implementing a contention-aware scheduler in the Linux kernel. Compared to the default Linux scheduler, our contention-aware scheduler significantly reduces tail latency for latency-critical workloads (e.g., database, key-value store) by up to 36.09 %, and improves the overall system throughput by up to1.072×1.072 \times. Finally, using real-world cluster traces from Alibaba and Google, we demonstrate that SMTcheck incurs negligible profiling overheads (≈0.113%\approx 0.113 \%), making it practical for deployment in productionscale datacenter environments.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines