Shaving the Peaks: Taming Tail Latency for Managed Workloads via Disaggregated Garbage Collection
Hongtao Lyu, Yuhan Li, Mingyu Wu
摘要
Language runtimes are essential systems commonly used in multi-tenant cloud scenarios, such as interactive web services and other cloud workloads. They usually provide memory management services, or garbage collection (GC), to automatically reclaim memory and reduce the labor work of application developers. Recent concurrent collectors allow GC to co-run with application threads (mutators), which reduces application pauses and intends to improve the applications' tail latency. However, this work observes that periodic GC workloads remain a primary source of long tail latency, particularly in resource-constrained multi-tenant environments. In such settings, GC threads consume significant CPU resources, leading to severe performance contention with mutators.
To resolve the contention, this work presents DGC, a disaggregated GC architecture that exposes GC as an external service. DGC decouples the most costly marking phase in concurrent GC and offloads it to a disaggregated marking engine. Through a co-design of the GC marking algorithm and an RDMA-based software paging mechanism, DGC's disaggregated marking engine achieves performance on par with local execution while offloading marking to a remote node. To improve resource utilization, DGC introduces a global GC orchestrator to serve multiple runtimes while minimizing the conflicts due to the overlapping of individual GC triggering points. DGC is implemented on the OpenJDK HotSpot Java virtual machine, and the evaluation results on representative latency-sensitive applications show that DGC reduces P99 latency by up to 64.4% under moderate workloads and improves the peak goodput by up to 24.0%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- AIFM: High-Performance, Application-Integrated Far MemoryZhenyuan Ruan, Malte Schwarzkopf, Marcos K. Aguilera, Adam BelayOSDI 2020 · 被引用 224 次
- Can far memory improve job throughput?Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout 等EuroSys 2020 · 被引用 163 次
- Empowering Azure Storage with RDMAWei Bai, Shanim Sainul Abdeen, Ankit Agrawal, Krishan Kumar Attre 等NSDI 2023 · 被引用 117 次
- LineFS: Efficient SmartNIC Offload of a Distributed File System with Pipeline ParallelismJongyul Kim, Insu Jang, Waleed Reda, Jaeseong Im 等SOSP 2021 · 被引用 83 次
- Characterizing Off-path SmartNIC for Accelerating Distributed SystemsXingda Wei, Rongxin Cheng, Yuhan Yang, Rong Chen 等OSDI 2023 · 被引用 68 次
相关 Paper
- Let It Go: Relieving Garbage Collection Pain for Latency Critical Applications in GolangJunxian Zhao, Xiaobo Zhou, Sang-Yoon Chang, Chengzhong XuHPDC 2023
- Semeru: A Memory-Disaggregated Managed RuntimeChenxi Wang, Haoran Ma, Shi Liu, Yuanqi Li 等OSDI 2020 · 被引用 6 次
- Mako: a low-pause, high-throughput evacuating collector for memory-disaggregated datacentersHaoran Ma, Shi Liu, Chenxi Wang, Yifan Qiao 等PLDI 2022 · 被引用 19 次
- Platinum: A CPU-Efficient Concurrent Garbage Collector for Tail-Reduction of Interactive ServicesMingyu Wu, Ziming Zhao, Yanfei Yang, Haoyu Li 等USENIX ATC 2020 · 被引用 18 次
- Jade: A High-throughput Concurrent Copying Garbage CollectorMingyu Wu, Liang Mao, Yude Lin, Yifeng Jin 等EuroSys 2024 · 被引用 5 次
