Criticality-Aware Instruction-Centric Bandwidth Partitioning for Data Center Applications
Liren Zhu, Liujia Li, Jianyu Wu, Yiming Yao, Zhan Shi, Jie Zhang, Zhenlin Wang, Xiaolin Wang, Yingwei Luo, Diyu Zhou
Abstract
To reduce operational costs, modern data centers co-locate high-priority latency-critical (LC) tasks and low-priority best-effort (BE) tasks on the same physical node to increase resource utilization. However, such co-location leads to contention for memory bandwidth, resulting in priority inversion, where BE tasks severely slow down LC tasks. This priority inversion often leads to violations of the quality of service (QoS) requirements for LC tasks, defeating the purpose of co-location. Prior approaches to this issue either fail to enforce the QoS requirements for LC tasks or underutilize memory bandwidth.We present Pivot, a novel bandwidth partitioning system that overcomes the limitations of prior approaches based on two key insights. First, memory accesses from LC tasks must be prioritized across all the components on the memory path rather than a single component, as done in prior work. Second, only the scheduling of a selective portion of performance-critical loads (i.e., those causing a long stall on the re-order buffer), instead of all memory accesses from LC tasks, should be prioritized. To leverage these insights, Pivot overcomes the key challenge of accurately identifying performance-critical loads while incurring minimal runtime overhead by proposing a two-phase profiling technique. Our extensive evaluation shows that Pivot improves effective machine utilization by up to while increasing the throughput of the BE applications by up to compared to state-of-the-art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 213 citations
- CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale ComputersTirthak Patel, Devesh TiwariHPCA 2020 · 153 citations
- Hermes: Accelerating Long-Latency Load Requests via Perceptron-Based Off-Chip Load PredictionRahul Bera, Konstantinos Kanellopoulos, Shankar Balachandran, David Novo et al.MICRO 2022 · 37 citations
- Ripple: Profile-Guided Instruction Cache Replacement for Data Center ApplicationsTanvir Ahmed Khan, Dexin Zhang, Akshitha Sriraman, Joseph Devietti et al.ISCA 2021 · 33 citations
- CRISP: critical slice prefetchingHeiner Litz, Grant Ayers, Parthasarathy RanganathanASPLOS 2022 · 33 citations
Related papers
- OLPart: Online Learning based Resource Partitioning for Colocating Multiple Latency-Critical Jobs on Commodity ComputersRuobing Chen, Haosen Shi, Yusen Li, Xiaoguang Liu et al.EuroSys 2023 · 27 citations
- Rhythm: component-distinguishable workload deployment in datacentersLaiping Zhao, Yanan Yang, Kaixuan Zhang, Xiaobo Zhou et al.EuroSys 2020 · 49 citations
- Holmes: SMT Interference Diagnosis and CPU Scheduling for Job Co-locationAidi Pi, Xiaobo Zhou, Chengzhong XuHPDC 2022 · 11 citations
- UFO: The Ultimate QoS-Aware Core Management for Virtualized and Oversubscribed Public CloudsYajuan Peng, Shuang Chen, Yi Zhao, Zhibin YuNSDI 2024 · 7 citations
- Ah-Q: Quantifying and Handling the Interference within a Datacenter from a System PerspectiveYuhang Liu, Xin Deng, Jiapeng Zhou, Mingyu Chen et al.HPCA 2023 · 15 citations
