USENIX Security2023Top-tier venue
Systematic Assessment of Fuzzers using Mutation Analysis
Philipp Görz, Björn Mathis, Keno Hassler, Emre Güler, Thorsten Holz, Andreas Zeller, Rahul Gopinath
Abstract
Fuzzing is an important method to discover vulnerabilities in programs. Despite considerable progress in this area in the past years, measuring and comparing the effectiveness of fuzzers is still an open research question. In software testing, the gold standard for evaluating test quality is mutation analysis, which evaluates a test's ability to detect synthetic bugs: If a set of tests fails to detect such mutations, it is expected to also fail to detect real bugs. Mutation analysis subsumes various coverage measures and provides a large and diverse set of faults that can be arbitrarily hard to trigger and detect, thus preventing the problems of saturation and overfitting. Unfortunately, the cost of traditional mutation analysis is exorbitant for fuzzing, as mutations need independent evaluation. In this paper, we apply modern mutation analysis techniques that pool multiple mutations and allow us -- for the first time -- to evaluate and compare fuzzers with mutation analysis. We introduce an evaluation bench for fuzzers and apply it to a number of popular fuzzers and subjects. In a comprehensive evaluation, we show how we can use it to assess fuzzer performance and measure the impact of improved techniques. The required CPU time remains manageable: 4.09 CPU years are needed to analyze a fuzzer on seven subjects and a total of 141,278 mutations. We find that today's fuzzers can detect only a small percentage of mutations, which should be seen as a challenge for future research -- notably in improving (1) detecting failures beyond generic crashes (2) triggering mutations (and thus faults).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac2168d9-8c26-4fba-af8b-30565af53428Cited by top-tier papers4
- Evaluating Directed Fuzzers: Are We Heading in the Right Direction?Tae Eun Kim, Jaeseung Choi, Seongjae Im, Kihong Heo et al.FSE 2024 · 7 citations
- pPatch: Automated Vulnerability UnpatchingTianyi Jing, Pengyu Ding, Meng Xu, Yinhao Hu et al.FSE 2026
- MBFuzzer: A Multi-Party Protocol Fuzzer for MQTT BrokersXiangpu Song, Jianliang Wu, Yingpei Zeng, Hao Pan et al.USENIX Security 2025
- Metamorphic CoverageJinsheng Ba, Yuancheng Jiang, Manuel RiggerISSTA 2026
Builds on9
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei et al.CCS 2018 · 753 citations
- REDQUEEN: Fuzzing with Input-to-State CorrespondenceCornelius Aschermann, Sergej Schumilo, Tim Blazytko, Robert Gawlik et al.NDSS 2019 · 413 citations
- LAVA: Large-Scale Automated Vulnerability AdditionBrendan Dolan-Gavitt, Patrick Hulin, Engin Kirda, Tim Leek et al.S&P 2016 · 354 citations
- On the Reliability of Coverage-Based Fuzzer BenchmarkingMarcel Böhme, László Szekeres, Jonathan MetzmanICSE 2022 · 91 citations
- Revisiting the Relationship Between Fault Detection, Test Adequacy Criteria, and Test Set SizeYiqun T. Chen, Rahul Gopinath, Anita Tadakamalla, Michael D. Ernst et al.ASE 2020 · 48 citations
Related papers
- Seed selection for successful fuzzingAdrian Herrera, Hendra Gunadi, Shane Magrath, Michael Norrish et al.ISSTA 2021 · 95 citations
- DARWIN: Survival of the Fittest Fuzzing MutatorsPatrick Jauernig, Domagoj Jakobovic, Stjepan Picek, Emmanuel Stapf et al.NDSS 2023
- Guiding Greybox Fuzzing with Mutation TestingVasudev Vikram, Isabella Laybourn, Ao Li, Nicole Nair et al.ISSTA 2023 · 22 citations
- On the use of mutation analysis for evaluating student test suite qualityJames Perretta, Andrew DeOrio, Arjun Guha, Jonathan BellISSTA 2022 · 6 citations
- Program Feature-Based Benchmarking for Fuzz TestingMiao Miao, Sriteja Kummita, Eric Bodden, Shiyi WeiISSTA 2025
