Harvesting Memory-bound CPU Stall Cycles in Software with MSH
Zhihong Luo, Sam Son, Sylvia Ratnasamy, Scott Shenker
摘要
Memory-bound stalls account for a significant portion of CPU cycles in datacenter workloads, which makes harvesting them to execute other useful work highly valuable. However, mainstream implementations of the hardware harvesting mechanism, simultaneous multithreading (SMT), are unsatisfactory. They incur high latency overhead and do not offer fine-grained configurability of the trade-off between latency and harvesting throughput, which hinders wide adoption for latency-critical services; and they support only limited degrees of concurrency, which prevents full harvesting of memory stall cycles.
We present MSH, the first system that transparently and efficiently harvests memory-bound stall cycles in software. MSH makes full use of stall cycles with concurrency scaling, while incurring minimal and configurable latency overhead. MSH achieves these with a novel co-design of profiling, program analysis, binary instrumentation and runtime scheduling. Our evaluation shows that MSH achieves up to 72% harvesting throughput of SMT for latency SLOs under which SMT has to be disabled, and that strategically combining MSH with SMT leads to higher throughput than SMT due to MSH's capability to fully harvest memory-bound stall cycles.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Understanding and Profiling CXL.mem Using PathFinderXiao Li, Zerui Guo, Yuebin Bai, Mahesh Ketkar 等SIGCOMM 2025 · 被引用 6 次
- AUM: Unleashing the Efficiency Potential of Shared Processors with Accelerator Units for LLM ServingXinkai Wang, Chao Li, Yiming Zhuansun, Jinyang Guo 等HPCA 2026 · 被引用 2 次
- Harvesting Spare CPU Resources in Container SystemsAdam Hall, Anirudh Sarma, Esha Choukse, Umakishore Ramachandran 等NSDI 2026 · 被引用 2 次
- Performance Predictability in Heterogeneous MemoryJinshu Liu, Hanchen Xu, Daniel S. Berger, Marcos K. Aguilera 等ASPLOS 2026 · 被引用 1 次
- Tierce: Observability-Driven Tiered Memory Management for Colocated WorkloadsHanchen Xu, Berkay Inceisci, Hao Li, Zhenyu Zhang 等SOSP 2026
它引用的顶会 Paper22
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 被引用 213 次
- CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale ComputersTirthak Patel, Devesh TiwariHPCA 2020 · 被引用 153 次
- Providing SLOs for Resource-Harvesting VMs in Cloud PlatformsPradeep Ambati, Iñigo Goiri, Felipe Vieira Frujeri, Alper Gun 等OSDI 2020 · 被引用 101 次
- RedLeaf: Isolation and Communication in a Safe Operating SystemVikram Narayanan, Tianjiao Huang, David Detweiler, Dan Appel 等OSDI 2020 · 被引用 86 次
- Classifying Memory Access Patterns for PrefetchingGrant Ayers, Heiner Litz, Christos Kozyrakis, Parthasarathy RanganathanASPLOS 2020 · 被引用 83 次
相关 Paper
- Harvesting Sub-Microsecond CXL Memory Stalls with LiteSwitchNanqinqin Li, Yuhong Zhong, Asaf Cidon, Michael J. FreedmanOSDI 2026
- Ghost Threading: Helper-Thread Prefetching for Real SystemsYuxin Guo, Akshay Bhosale, Utpal Bora, Alexandra W. Chadwick 等MICRO 2025 · 被引用 2 次
- HardHarvest: Hardware-Supported Core Harvesting for MicroservicesJovan Stojkovic, Chunao Liu, Muhammad Shahbaz, Josep TorrellasISCA 2025 · 被引用 4 次
- Holmes: SMT Interference Diagnosis and CPU Scheduling for Job Co-locationAidi Pi, Xiaobo Zhou, Chengzhong XuHPDC 2022 · 被引用 11 次
- SMTcheck: Accurate SMT Interference Prediction to Improve Scheduling Efficiency in DatacentersSanghyun Kim, Jinhyeok Oh, Taehun Kim, Gyutae Kim 等HPCA 2026
