NaturalFuzz: Natural Input Generation for Big Data Analytics
Ahmad Humayun, Yaoxuan Wu, Miryung Kim, Muhammad Ali Gulzar
摘要
Fuzzing applies input mutations iteratively with the only goal of finding more bugs, resulting in synthetic tests that tend to lack realism. Big data analytics are expected to ingest real-world data as input. Therefore, when synthetic test data are not easily comprehensible, they are less likely to facilitate the downstream task of fixing errors. Our position is that fuzzing in this domain must achieve both high naturalness and high code coverage. We propose a new natural synthetic test generation tool for big data analytics, called NATURALFUZZ. It generates both unstructured, semi-structured, and structured data with corresponding semantics such as 'zipcode' and 'age.' The key insights behind NATURALFUZZ are two-fold. First, though existing test data may be small and lack coverage, we can grow this data to increase code coverage. Second, we can strategically mix constituent parts across different rows and columns to construct new realistic synthetic data by leveraging fine-grained data provenance. On commercial big data application benchmarks, NATU-RALFUZZ achieves an additional 19.9% coverage and detects 1.9× more faults than a machine learning-based synthetic data generator (SDV) when generating comparably sized inputs. This is because an ML-based synthetic data generator does not consider which code branches are exercised by which input rows from which tables, while NATURALFUZZ is able to select input rows that have a high potential to increase code coverage and mutate the selected data towards unseen, new program behavior. NATURALFUZZ's test data is more realistic than the test data generated by two baseline fuzzers (BigFuzz and Jazzer), while increasing code coverage and fault detection potential. NATURALFUZZ is the first fuzzing methodology with three benefits: (1) exclusively generate natural inputs, (2) fuzz multiple input sources simultaneously, and (3) find deeper semantics faults.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SoK: Prudent Evaluation Practices for FuzzingMoritz Schloegel, Nils Bars, Nico Schiller, Lukas Bernhard 等S&P 2024 · 被引用 69 次
- Natural Symbolic Execution-Based Testing for Big Data AnalyticsYaoxuan Wu, Ahmad Humayun, Muhammad Ali Gulzar, Miryung KimFSE 2024 · 被引用 5 次
它引用的顶会 Paper6
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 被引用 1,026 次
- T-Fuzz: Fuzzing by Program TransformationHui Peng, Yan Shoshitaishvili, Mathias PayerS&P 2018 · 被引用 326 次
- PATA: Fuzzing with Path Aware Taint AnalysisJie Liang, Mingzhe Wang, Chijin Zhou, Zhiyong Wu 等S&P 2022 · 被引用 84 次
- BigFuzz: Efficient Fuzz Testing for Data Analytics Using Framework AbstractionQian Zhang, Jiyuan Wang, Muhammad Ali Gulzar, Rohan Padhye 等ASE 2020 · 被引用 27 次
- TaintStream: fine-grained taint tracking for big data platforms through dynamic code translationChengxu Yang, Yuanchun Li, Mengwei Xu, Zhenpeng Chen 等FSE 2021 · 被引用 11 次
相关 Paper
- Co-dependence Aware Fuzzing for Dataflow-Based Big Data AnalyticsAhmad Humayun, Miryung Kim, Muhammad Ali GulzarFSE 2023 · 被引用 2 次
- Low-Cost and Comprehensive Non-textual Input Fuzzing with LLM-Synthesized Input GeneratorsKunpeng Zhang, Zongjie Li, Daoyuan Wu, Shuai Wang 等USENIX Security 2025
- SmartFuzz: Leveraging Large Language Models and Feature Composition to Generate High-Quality Seeds for Database FuzzingLi Lin, Jintai Hong, Yanlin Zhuang, Rongxin WuOOPSLA 2026
- Semantic Image Fuzzing of AI Perception SystemsTrey Woodlief, Sebastian G. Elbaum, Kevin SullivanICSE 2022 · 被引用 17 次
- Skyfire: Data-Driven Seed Generation for FuzzingJunjie Wang, Bihuan Chen, Lei Wei, Yang LiuS&P 2017 · 被引用 382 次
