Twig: Multi-Agent Task Management for Colocated Latency-Critical Cloud Services
Rajiv Nishtala, Vinicius Petrucci, Paul M. Carpenter, Magnus Själander
摘要
Many of the important services running on data centres are latency-critical, time-varying, and demand strict user satisfaction. Stringent tail-latency targets for colocated services and increasing system complexity make it challenging to reduce the power consumption of data centres. Data centres typically sacrifice server efficiency to maintain tail-latency targets resulting in an increased total cost of ownership.
This paper introduces Twig, a scalable quality-of-service (QoS) aware task manager for latency-critical services co-located on a server system. Twig successfully leverages deep reinforcement learning to characterise tail latency using hardware performance counters and to drive energy-efficient task management decisions in data centres. We evaluate Twig on a typical data centre server managing four widely used latency-critical services. Our results show that Twig outperforms prior works in reducing energy usage by up to 38% while achieving up to 99% QoS guarantee for latency-critical services.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the EdgeChunlin Tian, Xinpeng Qin, Kahou Tam, Li Li 等USENIX ATC 2025 · 被引用 41 次
- EcoFaaS: Rethinking the Design of Serverless Environments for Energy EfficiencyJovan Stojkovic, Nikoleta Iliakopoulou, Tianyin Xu, Hubertus Franke 等ISCA 2024 · 被引用 29 次
- Accelerating bandwidth-bound deep learning inference with main-memory acceleratorsBenjamin Y. Cho, Jeageun Jung, Mattan ErezSC 2021 · 被引用 24 次
- CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable MulticoresNeeraj Kulkarni, Gonzalo Gonzalez-Pumariega, Amulya Khurana, Christine A. Shoemaker 等MICRO 2020 · 被引用 21 次
相关 Paper
- ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the CloudShuang Chen, Angela Jin, Christina Delimitrou, José F. MartínezHPCA 2022 · 被引用 34 次
- EcoCore: Dynamic Core Management for Improving Energy Efficiency in Latency-Critical ApplicationsGyeongseo Park, Minho Kim, Ki-Dong Kang, Yunhyeong Jeon 等MICRO 2025 · 被引用 1 次
- NMAP: Power Management Based on Network Packet Processing Mode Transition for Latency-Critical WorkloadsKi-Dong Kang, Gyeongseo Park, Hyosang Kim, Mohammad Alian 等MICRO 2021 · 被引用 18 次
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang 等ISCA 2020 · 被引用 149 次
- Criticality-Aware Instruction-Centric Bandwidth Partitioning for Data Center ApplicationsLiren Zhu, Liujia Li, Jianyu Wu, Yiming Yao 等HPCA 2025 · 被引用 4 次
