Co-dependence Aware Fuzzing for Dataflow-Based Big Data Analytics
Ahmad Humayun, Miryung Kim, Muhammad Ali Gulzar
摘要
Data-intensive scalable computing has become popular due to the increasing demands of analyzing big data. For example, Apache Spark and Hadoop allow developers to write dataflow-based applications with user-defined functions to process data with custom logic. Testing such applications is difficult. (1) These applications often take multiple datasets as input. (2) Unlike in SQL, there is no explicit schema for these datasets and each unstructured (or semi-structured) dataset is segmented and parsed at runtime. (3) Dataflow operators (e.g., join) create implicit co-dependence constraints between the fields of multiple datasets. An efficient and effective testing technique must analyze co-dependence among different regions of multiple datasets at the level of rows and columns and orchestrate input mutations jointly on co-dependent regions.
We propose DepFuzz to increase the effectiveness and efficiency of fuzz testing dataflow-based big data applications. The key insight behind DepFuzz is twofold. It keeps track of which code segments operate on which datasets, which rows, and which columns. By analyzing the use of dataflow operators (e.g., join and groupByKey) in tandem with the semantics of UDFs, DepFuzz generates test data that subsequently reach hard-to-reach regions of the application code. In real-world big data applications, DepFuzz finds 3.4× more faults, achieving 29% more statement coverage in half the time as Jazzer's, a state-of-the-art commercial fuzzer for Java bytecode. It outperforms prior DISC testing by exposing deeper semantic faults beyond simpler input formatting errors, especially when multiple datasets have complex interactions through dataflow operators.
• Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 被引用 1,026 次
- Driller: Augmenting Fuzzing Through Selective Symbolic ExecutionNick Stephens, John Grosen, Christopher Salls, Andrew Dutcher 等NDSS 2016 · 被引用 1,021 次
- LAVA: Large-Scale Automated Vulnerability AdditionBrendan Dolan-Gavitt, Patrick Hulin, Engin Kirda, Tim Leek 等S&P 2016 · 被引用 354 次
- T-Fuzz: Fuzzing by Program TransformationHui Peng, Yan Shoshitaishvili, Mathias PayerS&P 2018 · 被引用 326 次
- PATA: Fuzzing with Path Aware Taint AnalysisJie Liang, Mingzhe Wang, Chijin Zhou, Zhiyong Wu 等S&P 2022 · 被引用 84 次
相关 Paper
- BigFuzz: Efficient Fuzz Testing for Data Analytics Using Framework AbstractionQian Zhang, Jiyuan Wang, Muhammad Ali Gulzar, Rohan Padhye 等ASE 2020 · 被引用 27 次
- ECFuzz: Effective Configuration Fuzzing for Large-Scale SystemsJunqiang Li, Senyi Li, Keyao Li, Falin Luo 等ICSE 2024 · 被引用 12 次
- DSFuzz: Detecting Deep State Bugs with Dependent State ExplorationYinxi Liu, Wei MengCCS 2023 · 被引用 4 次
- FuzzyFlow: Leveraging Dataflow To Find and Squash Program Optimization BugsPhilipp Schaad, Timo Schneider, Tal Ben-Nun, Alexandru Calotoiu 等SC 2023 · 被引用 3 次
- MUZZ: Thread-aware Grey-box Fuzzing for Effective Bug Hunting in Multithreaded ProgramsHongxu Chen, Shengjian Guo, Yinxing Xue, Yulei Sui 等USENIX Security 2020
