Co-dependence Aware Fuzzing for Dataflow-Based Big Data Analytics
Ahmad Humayun, Miryung Kim, Muhammad Ali Gulzar
Abstract
Data-intensive scalable computing has become popular due to the increasing demands of analyzing big data. For example, Apache Spark and Hadoop allow developers to write dataflow-based applications with user-defined functions to process data with custom logic. Testing such applications is difficult. (1) These applications often take multiple datasets as input. (2) Unlike in SQL, there is no explicit schema for these datasets and each unstructured (or semi-structured) dataset is segmented and parsed at runtime. (3) Dataflow operators (e.g., join) create implicit co-dependence constraints between the fields of multiple datasets. An efficient and effective testing technique must analyze co-dependence among different regions of multiple datasets at the level of rows and columns and orchestrate input mutations jointly on co-dependent regions.
We propose DepFuzz to increase the effectiveness and efficiency of fuzz testing dataflow-based big data applications. The key insight behind DepFuzz is twofold. It keeps track of which code segments operate on which datasets, which rows, and which columns. By analyzing the use of dataflow operators (e.g., join and groupByKey) in tandem with the semantics of UDFs, DepFuzz generates test data that subsequently reach hard-to-reach regions of the application code. In real-world big data applications, DepFuzz finds 3.4× more faults, achieving 29% more statement coverage in half the time as Jazzer's, a state-of-the-art commercial fuzzer for Java bytecode. It outperforms prior DISC testing by exposing deeper semantic faults beyond simpler input formatting errors, especially when multiple datasets have complex interactions through dataflow operators.
• Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b1d8ff98-00ef-4f0a-8052-bcdee7cc4279Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 1,026 citations
- Driller: Augmenting Fuzzing Through Selective Symbolic ExecutionNick Stephens, John Grosen, Christopher Salls, Andrew Dutcher et al.NDSS 2016 · 1,021 citations
- LAVA: Large-Scale Automated Vulnerability AdditionBrendan Dolan-Gavitt, Patrick Hulin, Engin Kirda, Tim Leek et al.S&P 2016 · 354 citations
- T-Fuzz: Fuzzing by Program TransformationHui Peng, Yan Shoshitaishvili, Mathias PayerS&P 2018 · 326 citations
- PATA: Fuzzing with Path Aware Taint AnalysisJie Liang, Mingzhe Wang, Chijin Zhou, Zhiyong Wu et al.S&P 2022 · 84 citations
Related papers
- BigFuzz: Efficient Fuzz Testing for Data Analytics Using Framework AbstractionQian Zhang, Jiyuan Wang, Muhammad Ali Gulzar, Rohan Padhye et al.ASE 2020 · 27 citations
- ECFuzz: Effective Configuration Fuzzing for Large-Scale SystemsJunqiang Li, Senyi Li, Keyao Li, Falin Luo et al.ICSE 2024 · 12 citations
- DSFuzz: Detecting Deep State Bugs with Dependent State ExplorationYinxi Liu, Wei MengCCS 2023 · 4 citations
- FuzzyFlow: Leveraging Dataflow To Find and Squash Program Optimization BugsPhilipp Schaad, Timo Schneider, Tal Ben-Nun, Alexandru Calotoiu et al.SC 2023 · 3 citations
- MUZZ: Thread-aware Grey-box Fuzzing for Effective Bug Hunting in Multithreaded ProgramsHongxu Chen, Shengjian Guo, Yinxing Xue, Yulei Sui et al.USENIX Security 2020
