RESIN: A Holistic Service for Dealing with Memory Leaks in Production Cloud Infrastructure
Chang Lou, Cong Chen, Peng Huang, Yingnong Dang, Si Qin, Xinsheng Yang, Xukun Li, Qingwei Lin, Murali Chintalapati
Abstract
Memory leak is a notorious issue. Despite the extensive efforts, addressing memory leaks in large production cloud systems remains challenging. Existing solutions incur high overhead and/or suffer from high inaccuracies.
This paper presents RESIN, a solution designed to holistically address memory leaks in production cloud infrastructure. RESIN takes a divide-and-conquer approach to tackle the challenges. It performs a low-overhead detection first with a robust bucketization-based pivot scheme to identify suspicious leaking entities. It then takes live heap snapshots at appropriate time points in carefully sampled leak entities. RESIN analyzes the collected snapshots for leak diagnosis. Finally, RESIN automatically mitigates detected leaks.
RESIN has been running in production in Microsoft Azure for 3 years. It reports on average 24 leak tickets each month with high accuracy and low overhead, and provides effective diagnosis reports. Its results translate into a 41× reduction of VM reboots caused by low memory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a4f2886-8c24-4e3f-9d7d-9444a9f27008Cited by top-tier papers8
- Automatic Root Cause Analysis via Large Language Models for Cloud IncidentsYinfang Chen, Huaibing Xie, Minghua Ma, Yu Kang et al.EuroSys 2024 · 175 citations
- Fail through the Cracks: Cross-System Interaction Failures in Modern Cloud SystemsLilia Tang, Chaitanya Bhandari, Yongle Zhang, Anna Karanika et al.EuroSys 2023 · 16 citations
- Vicious Cycles in Distributed Software SystemsShangshu Qian, Wen Fan, Lin Tan, Yongle ZhangASE 2023 · 5 citations
- EXIST: Enabling Extremely Efficient Intra-Service Tracing Observability in DatacentersXinkai Wang, Xiaofeng Hou, Chao Li, Yuancheng Li et al.ASPLOS 2025 · 4 citations
- Observability Is Eating Your Cores: Fine-Grained Analysis of Microservice Metrics with IPU-Hosted SketchesAlessandro Cornacchia, Theophilus A. Benson, Muhammad Bilal, Marco CaniniNSDI 2026 · 3 citations
Builds on6
- Understanding, Detecting and Localizing Partial Failures in Large System SoftwareChang Lou, Peng Huang, Scott SmithNSDI 2020 · 88 citations
- Gandalf: An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud InfrastructureZe Li, Qian Cheng, Ken Hsieh, Yingnong Dang et al.NSDI 2020 · 69 citations
- Understanding and Detecting Software Upgrade Failures in Distributed SystemsYongle Zhang, Junwen Yang, Zhuqi Jin, Utsav Sethi et al.SOSP 2021 · 40 citations
- Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM InterruptionsSebastien Levy, Randolph Yao, Youjiang Wu, Yingnong Dang et al.OSDI 2020 · 35 citations
- MLEE: Effective Detection of Memory Leaks on Early-Exit Paths in OS KernelsWenwen WangUSENIX ATC 2021 · 15 citations
Related papers
- Amur: Fixing Multi-Resource Leaks Guided by Resource Flow AnalysisJinyoung Kim, Eunseok LeeASE 2025
- From Leaks to Fixes: Automated Repairs for Resource Leak WarningsAkshay Utture, Jens PalsbergFSE 2023 · 5 citations
- Project-Level Resource Leak Detection through Agent-based Ownership Analysis and Repair Pattern VerificationChengxin Xu, Xiu Zhang, Xiaorui GongICSE 2026
- JLeaks: A Featured Resource Leak Repository Collected From Hundreds of Open-Source Java ProjectsTianyang Liu, Weixing Ji, Xiaohui Dong, Wuhuang Yao et al.ICSE 2024 · 1 citation
- Memory-harvesting VMs in cloud platformsAlexander Fuerst, Stanko Novakovic, Iñigo Goiri, Gohar Irfan Chaudhry et al.ASPLOS 2022 · 39 citations
