Skyfire: Data-Driven Seed Generation for Fuzzing
Junjie Wang, Bihuan Chen, Lei Wei, Yang Liu
Abstract
Programs that take highly-structured files as inputs normally process inputs in stages: syntax parsing, semantic checking, and application execution. Deep bugs are often hidden in the application execution stage, and it is non-trivial to automatically generate test inputs to trigger them. Mutation-based fuzzing generates test inputs by modifying well-formed seed inputs randomly or heuristically. Most inputs are rejected at the early syntax parsing stage. Differently, generation-based fuzzing generates inputs from a specification (e.g., grammar). They can quickly carry the fuzzing beyond the syntax parsing stage. However, most inputs fail to pass the semantic checking (e.g., violating semantic rules), which restricts their capability of discovering deep bugs. In this paper, we propose a novel data-driven seed generation approach, named Skyfire, which leverages the knowledge in the vast amount of existing samples to generate well-distributed seed inputs for fuzzing programs that process highly-structured inputs. Skyfire takes as inputs a corpus and a grammar, and consists of two steps. The first step of Skyfire learns a probabilistic contextsensitive grammar (PCSG) to specify both syntax features and semantic rules, and then the second step leverages the learned PCSG to generate seed inputs. We fed the collected samples and the inputs generated by Skyfire as seeds of AFL to fuzz several open-source XSLT and XML engines (i.e., Sablotron, libxslt, and libxml2). The results have demonstrated that Skyfire can generate well-distributed inputs and thus significantly improve the code coverage (i.e., 20% for line coverage and 15% for function coverage on average) and the bug-finding capability of fuzzers. We also used the inputs generated by Skyfire to fuzz the closed-source JavaScript and rendering engine of Internet Explorer 11. Altogether, we discovered 19 new memory corruption bugs (among which there are 16 new vulnerabilities) and 32 denial-of-service bugs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3c1f133c-7ba6-4328-b497-2f93e5699b93Cited by top-tier papers104
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei et al.CCS 2018 · 753 citations
- CollAFL: Path Sensitive FuzzingShuitao Gan, Chao Zhang, Xiaojun Qin, Xuwen Tu et al.S&P 2018 · 426 citations
- T-Fuzz: Fuzzing by Program TransformationHui Peng, Yan Shoshitaishvili, Mathias PayerS&P 2018 · 326 citations
- NAUTILUS: Fishing for Deep Bugs with GrammarsCornelius Aschermann, Tommaso Frassetto, Thorsten Holz, Patrick Jauernig et al.NDSS 2019 · 291 citations
- Learning to Fuzz from Symbolic Execution with Application to Smart ContractsJingxuan He, Mislav Balunovic, Nodar Ambroladze, Petar Tsankov et al.CCS 2019 · 288 citations
Builds on2
Related papers
- FREEDOM: Engineering a State-of-the-Art DOM FuzzerWen Xu, Soyeon Park, Taesoo KimCCS 2020 · 25 citations
- SQUIRREL: Testing Database Management Systems with Language Validity and Coverage FeedbackRui Zhong, Yongheng Chen, Hong Hu, Hangfan Zhang et al.CCS 2020 · 5 citations
- SoFi: Reflection-Augmented Fuzzing for JavaScript EnginesXiaoyu He, Xiaofei Xie, Yuekang Li, Jianwen Sun et al.CCS 2021 · 34 citations
- Reducing Coverage-Equivalent Inputs in Grammar-Based Fuzzing by Avoiding Recurrent Rule SequencesJaehan Yoon, Yunji Seo, Hakjoo Oh, Sooyoung ChaFSE 2026
- Repair-Driven Greybox FuzzingBachir Bendrissou, Alastair F. Donaldson, Cristian CadarISSTA 2026
