Detecting cache-related bugs in Spark applications
Hui Li, Dong Wang, Tianze Huang, Yu Gao, Wensheng Dou, Lijie Xu, Wei Wang, Jun Wei, Hua Zhong
Abstract
Apache Spark has been widely used to build big data applications. Spark utilizes the abstraction of Resilient Distributed Dataset (RDD) to store and retrieve large-scale data. To reduce duplicate computation of an RDD, Spark can cache the RDD in memory and then reuse it later, thus improving performance. Spark relies on application developers to enforce caching decisions by using persist() and unpersist() APIs, e.g., which RDD is persisted and when the RDD is persisted / unpersisted. Incorrect RDD caching decisions can cause duplicate computations, or waste precious memory resource, thus introducing serious performance degradation in Spark applications. In this paper, we propose CacheCheck, to automatically detect cache-related bugs in Spark applications. We summarize six cache-related bug patterns in Spark applications, and then dynamically detect cache-related bugs by analyzing the execution traces of Spark applications. We evaluate CacheCheck on six real-world Spark applications. The experimental result shows that CacheCheck detects 72 previously unknown cache-related bugs, and 28 of them have been fixed by developers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Juggler: Autonomous Cost Optimization and Performance Prediction of Big Data ApplicationsHani Al-Sayeh, Bunjamin Memishi, Muhammad Attahir Jibril, Marcus Paradies et al.SIGMOD 2022 · 16 citations
- Agile-Ant: Self-managing Distributed Cache Management for Cost Optimization of Big Data ApplicationsHani Al-Sayeh, Muhammad Attahir Jibril, Kai-Uwe SattlerVLDB 2024 · 1 citation
Related papers
- Simulee: detecting CUDA synchronization bugs via memory-access modelingMingyuan Wu, Yicheng Ouyang, Husheng Zhou, Lingming Zhang et al.ICSE 2020 · 26 citations
- BigFuzz: Efficient Fuzz Testing for Data Analytics Using Framework AbstractionQian Zhang, Jiyuan Wang, Muhammad Ali Gulzar, Rohan Padhye et al.ASE 2020 · 27 citations
- DiffStream: differential output testing for stream processing programsKonstantinos Kallas, Filip Niksic, Caleb Stanford, Rajeev AlurOOPSLA 2020 · 18 citations
- AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency AnalysisXiang Fu, Weiping Zhang, Shiman Meng, Xin Huang et al.SC 2024 · 1 citation
- CrystalPerf: Learning to Characterize the Performance of Dataflow Computation through Code AnalysisHuangshi Tian, Minchen Yu, Wei WangUSENIX ATC 2021 · 2 citations
