BigFuzz: Efficient Fuzz Testing for Data Analytics Using Framework Abstraction
Qian Zhang, Jiyuan Wang, Muhammad Ali Gulzar, Rohan Padhye, Miryung Kim
摘要
As big data analytics become increasingly popular, data-intensive scalable computing (DISC) systems help address the scalability issue of handling large data. However, automated testing for such data-centric applications is challenging, because data is often incomplete, continuously evolving, and hard to know a priori. Fuzz testing has been proven to be highly effective in other domains such as security; however, it is nontrivial to apply such traditional fuzzing to big data analytics directly for three reasons: (1) the long latency of DISC systems prohibits the applicability of fuzzing: naïve fuzzing would spend 98% of the time in setting up a test environment; (2) conventional branch coverage is unlikely to scale to DISC applications because most binary code comes from the framework implementation such as Apache Spark; and (3) random bit or byte level mutations can hardly generate meaningful data, which fails to reveal real-world application bugs. We propose a novel coverage-guided fuzz testing tool for big data analytics, called BigFuzz. The key essence of our approach is that: (a) we focus on exercising application logic as opposed to increasing framework code coverage by abstracting the DISC framework using specifications. BigFuzz performs automated source to source transformations to construct an equivalent DISC application suitable for fast test generation, and (b) we design schema-aware data mutation operators based on our in-depth study of DISC application error types. BigFuzz speeds up the fuzzing time by 78 to 1477X compared to random fuzzing, improves application code coverage by 20% to 271%, and achieves 33% to 157% improvement in detecting application errors. When compared to the state of the art that uses symbolic execution to test big data analytics, BigFuzz is applicable to twice more programs and can find 81% more bugs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- QDiff: Differential Testing of Quantum Software StacksJiyuan Wang, Qian Zhang, Guoqing Harry Xu, Miryung KimASE 2021 · 被引用 48 次
- Subtle Bugs Everywhere: Generating Documentation for Data Wrangling CodeChenyang Yang, Shurui Zhou, Jin L. C. Guo, Christian KästnerASE 2021 · 被引用 25 次
- Randomized Testing of Byzantine Fault Tolerant AlgorithmsLevin N. Winter, Florena Buse, Daan de Graaf, Klaus von Gleissenthall 等OOPSLA 2023 · 被引用 23 次
- Guiding Greybox Fuzzing with Mutation TestingVasudev Vikram, Isabella Laybourn, Ao Li, Nicole Nair 等ISSTA 2023 · 被引用 22 次
- Understanding and Detecting On-The-Fly Configuration BugsTeng Wang, Zhouyang Jia, Shanshan Li, Si Zheng 等ICSE 2023 · 被引用 12 次
它引用的顶会 Paper6
- Driller: Augmenting Fuzzing Through Selective Symbolic ExecutionNick Stephens, John Grosen, Christopher Salls, Andrew Dutcher 等NDSS 2016 · 被引用 1,021 次
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei 等CCS 2018 · 被引用 753 次
- Full-Speed Fuzzing: Reducing Fuzzing Overhead through Coverage-Guided TracingStefan Nagy, Matthew HicksS&P 2019 · 被引用 156 次
- Debloating Software through Piece-Wise Compilation and LoadingAnh Quach, Aravind Prakash, Lok-Kwong YanUSENIX Security 2018 · 被引用 153 次
- RAZOR: A Framework for Post-deployment Software DebloatingChenxiong Qian, Hong Hu, Mansour Alharthi, Simon Pak Ho Chung 等USENIX Security 2019 · 被引用 132 次
相关 Paper
- NaturalFuzz: Natural Input Generation for Big Data AnalyticsAhmad Humayun, Yaoxuan Wu, Miryung Kim, Muhammad Ali GulzarASE 2023 · 被引用 2 次
- Co-dependence Aware Fuzzing for Dataflow-Based Big Data AnalyticsAhmad Humayun, Miryung Kim, Muhammad Ali GulzarFSE 2023 · 被引用 2 次
- T-Fuzz: Fuzzing by Program TransformationHui Peng, Yan Shoshitaishvili, Mathias PayerS&P 2018 · 被引用 326 次
- Prompt Fuzzing for Fuzz Driver GenerationYunlong Lyu, Yuxuan Xie, Peng Chen, Hao ChenCCS 2024 · 被引用 21 次
- Signal Breaker: Fuzzing Digital Signal ProcessorsCameron Santiago Garcia, Matthew HicksASPLOS 2026
