Ah-Q: Quantifying and Handling the Interference within a Datacenter from a System Perspective
Yuhang Liu, Xin Deng, Jiapeng Zhou, Mingyu Chen, Yungang Bao
摘要
Interference among applications frequently occurs in a datacenter and significantly influences the cost-efficiency and the user experience. However, it is challenging for us to quantify the exact intensity of the interference that occurred in the overall system of a datacenter, because there are many concurrent applications in a datacenter, and their type can be either latency-critical (LC) and best-effort (BE). To address this issue, we present the Ah-Q which includes a theory and a strategy.
First, we propose the "system entropy" (E S ) theory to holistically and analytically quantify the interference in a datacenter to address this vital issue. The interference is caused by the scarcity of resources or/and the irrationality of scheduling. As more appropriate scheduling can compensate for resource scarcity, we derive the concept of "resource equivalence" to quantify the effectiveness of a resource scheduling strategy. We evaluate different resource scheduling strategies to validate the correctness and effectiveness of the proposed theory.
Moreover, using the theory to eliminate interference, we propose a new resource scheduling strategy; i.e., ARQ, which dynamically allocates the isolated resources and the shared resources to simultaneously harvest the benefits of isolation and sharing. Our results show that compared to the state-of-theart strategies (PARTIES and CLITE), ARQ is more effective to reduce the tail latency of the LC applications and to increase the IPC of the BE applications. Compared with PARTIES and CLITE, ARQ increases the yield (the ratio of satisfactory LC applications) by 25% and 20%, respectively; when the load is low, ARQ increases IPC of BE applications by 63.8% and 37.1%, respectively; ARQ reduces E S by 36.4% and 33.3%, respectively. The effectiveness of ARQ has saved resources significantly to achieve the same satisfactory overall user experience.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- UFO: The Ultimate QoS-Aware Core Management for Virtualized and Oversubscribed Public CloudsYajuan Peng, Shuang Chen, Yi Zhao, Zhibin YuNSDI 2024 · 被引用 7 次
- Criticality-Aware Instruction-Centric Bandwidth Partitioning for Data Center ApplicationsLiren Zhu, Liujia Li, Jianyu Wu, Yiming Yao 等HPCA 2025 · 被引用 4 次
它引用的顶会 Paper7
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh 等ASPLOS 2021 · 被引用 226 次
- Caladan: Mitigating Interference at Microsecond TimescalesJoshua Fried, Zhenyuan Ruan, Amy Ousterhout, Adam BelayOSDI 2020 · 被引用 213 次
- CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale ComputersTirthak Patel, Devesh TiwariHPCA 2020 · 被引用 153 次
- Twig: Multi-Agent Task Management for Colocated Latency-Critical Cloud ServicesRajiv Nishtala, Vinicius Petrucci, Paul M. Carpenter, Magnus SjälanderHPCA 2020 · 被引用 76 次
- Rhythm: component-distinguishable workload deployment in datacentersLaiping Zhao, Yanan Yang, Kaixuan Zhang, Xiaobo Zhou 等EuroSys 2020 · 被引用 49 次
相关 Paper
- OLPart: Online Learning based Resource Partitioning for Colocating Multiple Latency-Critical Jobs on Commodity ComputersRuobing Chen, Haosen Shi, Yusen Li, Xiaoguang Liu 等EuroSys 2023 · 被引用 27 次
- INVAR: Inversion Aware Resource Provisioning and Workload Scheduling for Edge ComputingBin Wang, David Irwin, Prashant J. Shenoy, Don TowsleyINFOCOM 2024 · 被引用 11 次
- Intelligent Resource Scheduling for Co-located Latency-critical Services: A Multi-Model Collaborative Learning ApproachLei Liu, Xinglei Dou, Yuetao ChenFAST 2023
- Switches for HIRE: resource scheduling for data center in-network computingMarcel Blöcher, Lin Wang, Patrick Eugster, Max SchmidtASPLOS 2021 · 被引用 34 次
- Learnings from Deploying Network QoS Alignment to Application Priorities for Storage ServicesMatthew Buckley, Parsa Pazhooheshy, Z. Morley Mao, Nandita Dukkipati 等NSDI 2025 · 被引用 3 次
