Seed selection for successful fuzzing
Adrian Herrera, Hendra Gunadi, Shane Magrath, Michael Norrish, Mathias Payer, Antony L. Hosking
摘要
Mutation-based greybox fuzzing-unquestionably the most widelyused fuzzing technique-relies on a set of non-crashing seed inputs (a corpus) to bootstrap the bug-finding process. When evaluating a fuzzer, common approaches for constructing this corpus include: (i) using an empty file; (ii) using a single seed representative of the target's input format; or (iii) collecting a large number of seeds (e.g., by crawling the Internet). Little thought is given to how this seed choice affects the fuzzing process, and there is no consensus on which approach is best (or even if a best approach exists). To address this gap in knowledge, we systematically investigate and evaluate how seed selection affects a fuzzer's ability to find bugs in real-world software. This includes a systematic review of seed selection practices used in both evaluation and deployment contexts, and a large-scale empirical evaluation (over 33 CPU-years) of six seed selection approaches. These six seed selection approaches include three corpus minimization techniques (which select the smallest subset of seeds that trigger the same range of instrumentation data points as a full corpus). Our results demonstrate that fuzzing outcomes vary significantly depending on the initial seeds used to bootstrap the fuzzer, with minimized corpora outperforming singleton, empty, and large (in the order of thousands of files) seed sets. Consequently, we encourage seed selection to be foremost in mind when evaluating/deploying fuzzers, and recommend that (a) seed choice be carefully considered and explicitly documented, and (b) never to evaluate fuzzers with only a single seed. CCS CONCEPTS • Software and its engineering → Software testing and debugging; • Security and privacy → Software and application security.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper47
- On the Reliability of Coverage-Based Fuzzer BenchmarkingMarcel Böhme, László Szekeres, Jonathan MetzmanICSE 2022 · 被引用 91 次
- LLM-Fuzzer: Scaling Assessment of Large Language Model JailbreaksJiahao Yu, Xingwei Lin, Zheng Yu, Xinyu XingUSENIX Security 2024 · 被引用 83 次
- Effective Seed Scheduling for Fuzzing with Graph Centrality AnalysisDongdong She, Abhishek Shah, Suman JanaS&P 2022 · 被引用 78 次
- LibAFL: A Framework to Build Modular and Reusable FuzzersAndrea Fioraldi, Dominik Christian Maier, Dongjia Zhang, Davide BalzarottiCCS 2022 · 被引用 71 次
- SoK: Prudent Evaluation Practices for FuzzingMoritz Schloegel, Nils Bars, Nico Schiller, Lukas Bernhard 等S&P 2024 · 被引用 69 次
它引用的顶会 Paper22
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 被引用 1,026 次
- Directed Greybox FuzzingMarcel Böhme, Van-Thuan Pham, Manh-Dung Nguyen, Abhik RoychoudhuryCCS 2017 · 被引用 836 次
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei 等CCS 2018 · 被引用 753 次
- QSYM : A Practical Concolic Execution Engine Tailored for Hybrid FuzzingInsu Yun, Sangho Lee, Meng Xu, Yeongjin Jang 等USENIX Security 2018 · 被引用 537 次
- CollAFL: Path Sensitive FuzzingShuitao Gan, Chao Zhang, Xiaojun Qin, Xuwen Tu 等S&P 2018 · 被引用 426 次
相关 Paper
- An Empirical Examination of Fuzzer Mutator PerformanceJames Kukucka, Luís Pina, Paul Ammann, Jonathan BellISSTA 2024 · 被引用 4 次
- Systematic Assessment of Fuzzers using Mutation AnalysisPhilipp Görz, Björn Mathis, Keno Hassler, Emre Güler 等USENIX Security 2023
- RandSet: Randomized Corpus Reduction for Fuzzing Seed SchedulingYuchong Xie, Kaikai Zhang, Yu Liu, Rundong Yang 等OOPSLA 2026 · 被引用 1 次
- MendelFuzz: The Return of the Deterministic StageHan Zheng, Flavio Toffalini, Marcel Böhme, Mathias PayerFSE 2025 · 被引用 2 次
- Path Transitions Tell More: Optimizing Fuzzing Schedules via Runtime Program StatesKunpeng Zhang, Xi Xiao, Xiaogang Zhu, Ruoxi Sun 等ICSE 2022 · 被引用 25 次
