Low-latency, high-throughput garbage collection
Wenyu Zhao, Stephen M. Blackburn, Kathryn S. McKinley
Abstract
To achieve short pauses, state-of-the-art concurrent copying collectors such as C4, Shenandoah, and ZGC use substantially more CPU cycles and memory than simpler collectors. They suffer from design limitations: i) concurrent copying with inherently expensive read and write barriers, ii) scalability limitations due to tracing, and iii) immediacy limitations for mature objects that impose memory overheads.
This paper takes a different approach to optimizing responsiveness and throughput. It uses the insight that regular, brief stop-the-world collections deliver sufficient responsiveness at greater efficiency than concurrent evacuation. It introduces LXR, where stop-the-world collections use reference counting (RC) and judicious copying. RC delivers scalability and immediacy, promptly reclaiming young and mature objects. RC, in a hierarchical Immix heap structure, reclaims most memory without any copying. Occasional concurrent tracing identifies cyclic garbage. LXR introduces: i) RC remembered sets for judicious copying of mature objects; ii) a novel low-overhead write barrier that combines coalescing reference counting, concurrent tracing, and remembered set maintenance; iii) object reclamation while performing a concurrent trace; iv) lazy processing of decrements; and v) novel survival rate triggers that modulate pause durations.
LXR combines excellent responsiveness and throughput, improving over production collectors. On the widely-used Lucene search engine in a tight heap, LXR delivers 7.8× better throughput and 10× better 99.99% tail latency than Shenandoah. On 17 diverse modern workloads in a moderate heap, LXR outperforms OpenJDK's default G1 on throughput by 4% and Shenandoah by 43%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab83e4cc-a150-4433-99bb-98e4d577b876Cited by top-tier papers10
- More Apps, Faster Hot-Launch on Mobile Devices via Fore/Background-aware GC-Swap Co-designJiacheng Huang, Yunmo Zhang, Junqiao Qiu, Yu Liang et al.ASPLOS 2024 · 12 citations
- Concurrent Immediate Reference CountingJaehwang Jung, Jeonghyeon Kim, Matthew J. Parkinson, Jeehoon KangPLDI 2024 · 6 citations
- Jade: A High-throughput Concurrent Copying Garbage CollectorMingyu Wu, Liang Mao, Yude Lin, Yifeng Jin et al.EuroSys 2024 · 5 citations
- Work Packets: A New Abstraction for GC Software Engineering, Optimization, and InnovationWenyu Zhao, Stephen M. Blackburn, Kathryn S. McKinleyOOPSLA 2025 · 3 citations
- Evaluating Garbage Collection Performance Across Managed Language RuntimesYicheng Wang, Wensheng Dou, Yu Liang, Yi Wang et al.ICSE 2025 · 1 citation
Related papers
- Mark-Scavenge: Waiting for Trash to Take Itself OutJonas Norlinder, Erik Österlund, David Black-Schaffer, Tobias WrigstadOOPSLA 2024
- Advancing Performance via a Systematic Application of Research and Industrial Best PracticeWenyu Zhao, Stephen M. Blackburn, Kathryn S. McKinley, Man Cao et al.OOPSLA 2025
- FlexHeap: Dynamic I/O-Aware Heap Resizing for Managed ApplicationsIacovos G. Kolokasis, Shoaib Akram, Foivos S. Zakkak, Polyvios Pratikakis et al.PLDI 2026
- Mako: a low-pause, high-throughput evacuating collector for memory-disaggregated datacentersHaoran Ma, Shi Liu, Chenxi Wang, Yifan Qiao et al.PLDI 2022 · 19 citations
- Iso: Request-Private Garbage CollectionTianle Qiu, Stephen M. BlackburnPLDI 2025 · 1 citation
