Twig: Multi-Agent Task Management for Colocated Latency-Critical Cloud Services
Rajiv Nishtala, Vinicius Petrucci, Paul M. Carpenter, Magnus Själander
Abstract
Many of the important services running on data centres are latency-critical, time-varying, and demand strict user satisfaction. Stringent tail-latency targets for colocated services and increasing system complexity make it challenging to reduce the power consumption of data centres. Data centres typically sacrifice server efficiency to maintain tail-latency targets resulting in an increased total cost of ownership.
This paper introduces Twig, a scalable quality-of-service (QoS) aware task manager for latency-critical services co-located on a server system. Twig successfully leverages deep reinforcement learning to characterise tail latency using hardware performance counters and to drive energy-efficient task management decisions in data centres. We evaluate Twig on a typical data centre server managing four widely used latency-critical services. Our results show that Twig outperforms prior works in reducing energy usage by up to 38% while achieving up to 99% QoS guarantee for latency-critical services.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10b32ffc-0198-4ca2-a618-afbe5dc39f3eCited by top-tier papers13
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas et al.HPCA 2025 · 106 citations
- CLONE: Customizing LLMs for Efficient Latency-Aware Inference at the EdgeChunlin Tian, Xinpeng Qin, Kahou Tam, Li Li et al.USENIX ATC 2025 · 41 citations
- EcoFaaS: Rethinking the Design of Serverless Environments for Energy EfficiencyJovan Stojkovic, Nikoleta Iliakopoulou, Tianyin Xu, Hubertus Franke et al.ISCA 2024 · 29 citations
- Accelerating bandwidth-bound deep learning inference with main-memory acceleratorsBenjamin Y. Cho, Jeageun Jung, Mattan ErezSC 2021 · 24 citations
- CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable MulticoresNeeraj Kulkarni, Gonzalo Gonzalez-Pumariega, Amulya Khurana, Christine A. Shoemaker et al.MICRO 2020 · 21 citations
Related papers
- ReTail: Opting for Learning Simplicity to Enable QoS-Aware Power Management in the CloudShuang Chen, Angela Jin, Christina Delimitrou, José F. MartínezHPCA 2022 · 34 citations
- EcoCore: Dynamic Core Management for Improving Energy Efficiency in Latency-Critical ApplicationsGyeongseo Park, Minho Kim, Ki-Dong Kang, Yunhyeong Jeon et al.MICRO 2025 · 1 citation
- NMAP: Power Management Based on Network Packet Processing Mode Transition for Latency-Critical WorkloadsKi-Dong Kang, Gyeongseo Park, Hyosang Kim, Mohammad Alian et al.MICRO 2021 · 18 citations
- DeepRecSys: A System for Optimizing End-To-End At-Scale Neural Recommendation InferenceUdit Gupta, Samuel Hsia, Vikram Saraph, Xiaodong Wang et al.ISCA 2020 · 149 citations
- Criticality-Aware Instruction-Centric Bandwidth Partitioning for Data Center ApplicationsLiren Zhu, Liujia Li, Jianyu Wu, Yiming Yao et al.HPCA 2025 · 4 citations
