Harvesting Memory-bound CPU Stall Cycles in Software with MSH
Zhihong Luo, Sam Son, Sylvia Ratnasamy, Scott Shenker
Abstract
Memory-bound stalls account for a significant portion of CPU cycles in datacenter workloads, which makes harvesting them to execute other useful work highly valuable. However, mainstream implementations of the hardware harvesting mechanism, simultaneous multithreading (SMT), are unsatisfactory. They incur high latency overhead and do not offer fine-grained configurability of the trade-off between latency and harvesting throughput, which hinders wide adoption for latency-critical services; and they support only limited degrees of concurrency, which prevents full harvesting of memory stall cycles.
We present MSH, the first system that transparently and efficiently harvests memory-bound stall cycles in software. MSH makes full use of stall cycles with concurrency scaling, while incurring minimal and configurable latency overhead. MSH achieves these with a novel co-design of profiling, program analysis, binary instrumentation and runtime scheduling. Our evaluation shows that MSH achieves up to 72% harvesting throughput of SMT for latency SLOs under which SMT has to be disabled, and that strategically combining MSH with SMT leads to higher throughput than SMT due to MSH's capability to fully harvest memory-bound stall cycles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed9c16e4-5efe-4d31-8bb8-98ba645e8386Cited by top-tier papers7
- Understanding and Profiling CXL.mem Using PathFinderXiao Li, Zerui Guo, Yuebin Bai, Mahesh Ketkar et al.SIGCOMM 2025 · 6 citations
- AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM ServingXinkai Wang, Chao Li, Yiming Zhuansun, Jinyang Guo et al.HPCA 2026 · 2 citations
- Harvesting Spare CPU Resources in Container SystemsAdam Hall, Anirudh Sarma, Esha Choukse, Umakishore Ramachandran et al.NSDI 2026 · 2 citations
- Performance Predictability in Heterogeneous MemoryJinshu Liu, Hanchen Xu, Daniel S. Berger, Marcos K. Aguilera et al.ASPLOS 2026 · 1 citation
- Tierce: Observability-Driven Tiered Memory Management for Colocated WorkloadsHanchen Xu, Berkay Inceisci, Hao Li, Zhenyu Zhang et al.SOSP 2026
Builds on22
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale ComputersTirthak Patel, Devesh TiwariHPCA 2020 · 153 citations
- Providing SLOs for Resource-Harvesting VMs in Cloud PlatformsPradeep Ambati, Iñigo Goiri, Felipe Vieira Frujeri, Alper Gun et al.OSDI 2020 · 101 citations
- RedLeaf: Isolation and Communication in a Safe Operating SystemVikram Narayanan, Tianjiao Huang, David Detweiler, Dan Appel et al.OSDI 2020 · 86 citations
- Classifying Memory Access Patterns for PrefetchingGrant Ayers, Heiner Litz, Christos Kozyrakis, Parthasarathy RanganathanASPLOS 2020 · 83 citations
Related papers
- Harvesting Sub-Microsecond CXL Memory Stalls with LiteSwitchNanqinqin Li, Yuhong Zhong, Asaf Cidon, Michael J. FreedmanOSDI 2026
- Ghost Threading: Helper-Thread Prefetching for Real SystemsYuxin Guo, Akshay Bhosale, Utpal Bora, Alexandra W. Chadwick et al.MICRO 2025 · 2 citations
- HardHarvest: Hardware-Supported Core Harvesting for MicroservicesJovan Stojkovic, Chunao Liu, Muhammad Shahbaz, Josep TorrellasISCA 2025 · 4 citations
- Holmes: SMT Interference Diagnosis and CPU Scheduling for Job Co-locationAidi Pi, Xiaobo Zhou, Chengzhong XuHPDC 2022 · 11 citations
- SMTcheck: Accurate SMT Interference Prediction to Improve Scheduling Efficiency in DatacentersSanghyun Kim, Jinhyeok Oh, Taehun Kim, Gyutae Kim et al.HPCA 2026
