CuttleSys: Data-Driven Resource Management for Interactive Services on Reconfigurable Multicores
Neeraj Kulkarni, Gonzalo Gonzalez-Pumariega, Amulya Khurana, Christine A. Shoemaker, Christina Delimitrou, David H. Albonesi
摘要
Multi-tenancy for latency-critical applications leads to resource interference and unpredictable performance. Core reconfiguration opens up more opportunities for application colocation, as it allows the hardware to adjust to the dynamic performance and power needs of a specific mix of co-scheduled services. However, reconfigurability also introduces challenges, as even for a small number of reconfigurable cores, exploring the design space becomes more time- and resource-demanding.We present CuttleSys, a runtime for reconfigurable multicores that leverages scalable and lightweight data mining to quickly identify suitable core and cache configurations for a set of co-scheduled applications. The runtime combines collaborative filtering to infer the behavior of each job on every core and cache configuration, with Dynamically Dimensioned Search to efficiently explore the configuration space. We evaluate CuttleSys on multicores with tens of reconfigurable cores and show up to 2.46× and 1.55× performance improvements compared to core-level gating and oracle-like asymmetric multicores respectively, under stringent power constraints.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- VELTAIR: towards high-performance multi-tenant deep learning services via adaptive compilation and schedulingZihan Liu, Jingwen Leng, Zhihui Zhang, Quan Chen 等ASPLOS 2022 · 被引用 52 次
- Don't Forget the I/O When Allocating Your LLCYifan Yuan, Mohammad Alian, Yipeng Wang, Ren Wang 等ISCA 2021 · 被引用 37 次
- Ah-Q: Quantifying and Handling the Interference within a Datacenter from a System PerspectiveYuhang Liu, Xin Deng, Jiapeng Zhou, Mingyu Chen 等HPCA 2023 · 被引用 15 次
- ALTOCUMULUS: Scalable Scheduling for Nanosecond-Scale Remote Procedure CallsJiechen Zhao, Iris Uwizeyimana, Karthik Ganesan, Mark C. Jeffrey 等MICRO 2022 · 被引用 11 次
它引用的顶会 Paper1
相关 Paper
- CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale ComputersTirthak Patel, Devesh TiwariHPCA 2020 · 被引用 153 次
- Alita: comprehensive performance isolation through bias resource management for public cloudsQuan Chen, Shuai Xue, Shang Zhao, Shanpei Chen 等SC 2020 · 被引用 29 次
- SATORI: Efficient and Fair Resource Partitioning by Sacrificing Short-Term Benefits for Long-Term Gains*Rohan Basu Roy, Tirthak Patel, Devesh TiwariISCA 2021 · 被引用 34 次
- Co-Optimizing Cache Partitioning and Multi-Core Task Scheduling: Exploit Cache Sensitivity or Not?Binqi Sun, Debayan Roy, Tomasz Kloda, Andrea Bastoni 等RTSS 2023 · 被引用 4 次
- UFO: The Ultimate QoS-Aware Core Management for Virtualized and Oversubscribed Public CloudsYajuan Peng, Shuang Chen, Yi Zhao, Zhibin YuNSDI 2024 · 被引用 7 次
