SoK: Prudent Evaluation Practices for Fuzzing
Moritz Schloegel, Nils Bars, Nico Schiller, Lukas Bernhard, Tobias Scharnowski, Addison Crump, Arash Ale Ebrahim, Nicolai Bissantz, Marius Muench, Thorsten Holz
摘要
Fuzzing has proven to be a highly effective approach to uncover software bugs over the past decade. After AFL popularized the groundbreaking concept of lightweight coverage feedback, the field of fuzzing has seen a vast amount of scientific work proposing new techniques, improving methodological aspects of existing strategies, or porting existing methods to new domains. All such work must demonstrate its merit by showing its applicability to a problem, measuring its performance, and often showing its superiority over existing works in a thorough, empirical evaluation. Yet, fuzzing is highly sensitive to its target, environment, and circumstances, e. g., randomness in the testing process. After all, relying on randomness is one of the core principles of fuzzing, governing many aspects of a fuzzer’s behavior. Combined with the often highly difficult to control environment, the reproducibility of experiments is a crucial concern and requires a prudent evaluation setup. To address these threats to validity, several works, most notably Evaluating Fuzz Testing by Klees et al., have outlined how a carefully designed evaluation setup should be implemented, but it remains unknown to what extent their recommendations have been adopted in practice.In this work, we systematically analyze the evaluation of 150 fuzzing papers published at the top venues between 2018 and 2023. We study how existing guidelines are implemented and observe potential shortcomings and pitfalls. We find a surprising disregard of the existing guidelines regarding statistical tests and systematic errors in fuzzing evaluations. For example, when investigating reported bugs, we find that the search for vulnerabilities in real-world software leads to authors requesting and receiving CVEs of questionable quality. Extending our literature analysis to the practical domain, we attempt to reproduce claims of eight fuzzing papers. These case studies allow us to assess the practical reproducibility of fuzzing research and identify archetypal pitfalls in the evaluation design. Unfortunately, our reproduced results reveal several deficiencies in the studied papers, and we are unable to fully support and reproduce the respective claims. To help the field of fuzzing move toward a scientifically reproducible evaluation strategy, we propose updated guidelines for conducting a fuzzing evaluation that future work should follow.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper41
- SoK: The Pitfalls of Deep Reinforcement Learning for CybersecurityShae McFadden, Myles Foley, Elizabeth Bates, Ilias Tsingenopoulos 等USENIX Security 2026 · 被引用 7 次
- No Peer, no Cry: Network Application Fuzzing via Fault InjectionNils Bars, Moritz Schloegel, Nico Schiller, Lukas Bernhard 等CCS 2024 · 被引用 5 次
- CountDown: Refcount-guided Fuzzing for Exposing Temporal Memory Errors in Linux KernelShuangpeng Bai, Zhechang Zhang, Hong HuCCS 2024 · 被引用 4 次
- An Empirical Examination of Fuzzer Mutator PerformanceJames Kukucka, Luís Pina, Paul Ammann, Jonathan BellISSTA 2024 · 被引用 4 次
- DarthShader: Fuzzing WebGPU Shader Translators & CompilersLukas Bernhard, Nico Schiller, Moritz Schloegel, Nils Bars 等CCS 2024 · 被引用 3 次
它引用的顶会 Paper150
- Coverage-based Greybox Fuzzing as Markov ChainMarcel Böhme, Van-Thuan Pham, Abhik RoychoudhuryCCS 2016 · 被引用 1,026 次
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei 等CCS 2018 · 被引用 753 次
- Angora: Efficient Fuzzing by Principled SearchPeng Chen, Hao ChenS&P 2018 · 被引用 616 次
- QSYM : A Practical Concolic Execution Engine Tailored for Hybrid FuzzingInsu Yun, Sangho Lee, Meng Xu, Yeongjin Jang 等USENIX Security 2018 · 被引用 537 次
- CollAFL: Path Sensitive FuzzingShuitao Gan, Chao Zhang, Xiaojun Qin, Xuwen Tu 等S&P 2018 · 被引用 426 次
相关 Paper
- Seed selection for successful fuzzingAdrian Herrera, Hendra Gunadi, Shane Magrath, Michael Norrish 等ISSTA 2021 · 被引用 95 次
- The Use of Likely Invariants as Feedback for FuzzersAndrea Fioraldi, Daniele Cono D'Elia, Davide BalzarottiUSENIX Security 2021 · 被引用 67 次
- T-Fuzz: Fuzzing by Program TransformationHui Peng, Yan Shoshitaishvili, Mathias PayerS&P 2018 · 被引用 326 次
- UNIFUZZ: A Holistic and Pragmatic Metrics-Driven Platform for Evaluating FuzzersYuwei Li, Shouling Ji, Yuan Chen, Sizhuang Liang 等USENIX Security 2021 · 被引用 142 次
- Green Fuzzing: A Saturation-Based Stopping Criterion using Vulnerability PredictionStephan Lipp, Daniel Elsner, Severin Kacianka, Alexander Pretschner 等ISSTA 2023 · 被引用 6 次
