Detecting cache-related bugs in Spark applications
Hui Li, Dong Wang, Tianze Huang, Yu Gao, Wensheng Dou, Lijie Xu, Wei Wang, Jun Wei, Hua Zhong
摘要
Apache Spark has been widely used to build big data applications. Spark utilizes the abstraction of Resilient Distributed Dataset (RDD) to store and retrieve large-scale data. To reduce duplicate computation of an RDD, Spark can cache the RDD in memory and then reuse it later, thus improving performance. Spark relies on application developers to enforce caching decisions by using persist() and unpersist() APIs, e.g., which RDD is persisted and when the RDD is persisted / unpersisted. Incorrect RDD caching decisions can cause duplicate computations, or waste precious memory resource, thus introducing serious performance degradation in Spark applications. In this paper, we propose CacheCheck, to automatically detect cache-related bugs in Spark applications. We summarize six cache-related bug patterns in Spark applications, and then dynamically detect cache-related bugs by analyzing the execution traces of Spark applications. We evaluate CacheCheck on six real-world Spark applications. The experimental result shows that CacheCheck detects 72 previously unknown cache-related bugs, and 28 of them have been fixed by developers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Juggler: Autonomous Cost Optimization and Performance Prediction of Big Data ApplicationsHani Al-Sayeh, Bunjamin Memishi, Muhammad Attahir Jibril, Marcus Paradies 等SIGMOD 2022 · 被引用 16 次
- Agile-Ant: Self-managing Distributed Cache Management for Cost Optimization of Big Data ApplicationsHani Al-Sayeh, Muhammad Attahir Jibril, Kai-Uwe SattlerVLDB 2024 · 被引用 1 次
相关 Paper
- Simulee: detecting CUDA synchronization bugs via memory-access modelingMingyuan Wu, Yicheng Ouyang, Husheng Zhou, Lingming Zhang 等ICSE 2020 · 被引用 26 次
- BigFuzz: Efficient Fuzz Testing for Data Analytics Using Framework AbstractionQian Zhang, Jiyuan Wang, Muhammad Ali Gulzar, Rohan Padhye 等ASE 2020 · 被引用 27 次
- DiffStream: differential output testing for stream processing programsKonstantinos Kallas, Filip Niksic, Caleb Stanford, Rajeev AlurOOPSLA 2020 · 被引用 18 次
- AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency AnalysisXiang Fu, Weiping Zhang, Shiman Meng, Xin Huang 等SC 2024 · 被引用 1 次
- CrystalPerf: Learning to Characterize the Performance of Dataflow Computation through Code AnalysisHuangshi Tian, Minchen Yu, Wei WangUSENIX ATC 2021 · 被引用 2 次
