Learning input tokens for effective fuzzing
Björn Mathis, Rahul Gopinath, Andreas Zeller
摘要
Modern fuzzing tools like AFL operate at a lexical level: They explore the input space of tested programs one byte after another. For inputs with complex syntactical properties, this is very inefficient, as keywords and other tokens have to be composed one character at a time. Fuzzers thus allow to specify dictionaries listing possible tokens the input can be composed from; such dictionaries speed up fuzzers dramatically. Also, fuzzers make use of dynamic tainting to track input tokens and infer values that are expected in the input validation phase. Unfortunately, such tokens are usually implicitly converted to program specific values which causes a loss of the taints attached to the input data in the lexical phase. In this paper, we present a technique to extend dynamic tainting to not only track explicit data flows but also taint implicitly converted data without suffering from taint explosion. This extension makes it possible to augment existing techniques and automatically infer a set of tokens and seed inputs for the input language of a program given nothing but the source code. Specifically targeting the lexical analysis of an input processor, our lFuzzer test generator systematically explores branches of the lexical analysis, producing a set of tokens that fully cover all decisions seen. The resulting set of tokens can be directly used as a dictionary for fuzzing. Along with the token extraction seed inputs are generated which give further fuzzing processes a head start. In our experiments, the lFuzzer-AFL combination achieves up to 17% more coverage on complex input formats like json, lisp, tinyC, and JavaScript compared to AFL. CCS CONCEPTS • Software and its engineering → Software testing and debugging; Parsers; • Theory of computation → Regular languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- SMARTIAN: Enhancing Smart Contract Fuzzing with Static and Dynamic Data-Flow AnalysesJaeseung Choi, Doyeon Kim, Soomin Kim, Gustavo Grieco 等ASE 2021 · 被引用 164 次
- Automated conformance testing for JavaScript engines via deep compiler fuzzingGuixin Ye, Zhanyong Tang, Shin Hwei Tan, Songfang Huang 等PLDI 2021 · 被引用 75 次
- WhiteFox: White-Box Compiler Fuzzing Empowered by Large Language ModelsChenyuan Yang, Yinlin Deng, Runyu Lu, Jiayi Yao 等OOPSLA 2024 · 被引用 74 次
- The Use of Likely Invariants as Feedback for FuzzersAndrea Fioraldi, Daniele Cono D'Elia, Davide BalzarottiUSENIX Security 2021 · 被引用 67 次
- RULF: Rust Library Fuzzing via API Dependency Graph TraversalJianfeng Jiang, Hui Xu, Yangfan ZhouASE 2021 · 被引用 44 次
它引用的顶会 Paper4
- VUzzer: Application-aware Evolutionary FuzzingSanjay Rawat, Vivek Jain, Ashish Kumar, Lucian Cojocar 等NDSS 2017 · 被引用 700 次
- REDQUEEN: Fuzzing with Input-to-State CorrespondenceCornelius Aschermann, Sergej Schumilo, Tim Blazytko, Robert Gawlik 等NDSS 2019 · 被引用 413 次
- Skyfire: Data-Driven Seed Generation for FuzzingJunjie Wang, Bihuan Chen, Lei Wei, Yang LiuS&P 2017 · 被引用 382 次
- Mining input grammars from dynamic control flowRahul Gopinath, Björn Mathis, Andreas ZellerFSE 2020 · 被引用 60 次
相关 Paper
- Token-Level FuzzingChristopher Salls, Chani Jindal, Jake Corina, Christopher Kruegel 等USENIX Security 2021
- Low-Cost and Comprehensive Non-textual Input Fuzzing with LLM-Synthesized Input GeneratorsKunpeng Zhang, Zongjie Li, Daoyuan Wu, Shuai Wang 等USENIX Security 2025
- T-Fuzz: Fuzzing by Program TransformationHui Peng, Yan Shoshitaishvili, Mathias PayerS&P 2018 · 被引用 326 次
- Ijon: Exploring Deep State Spaces via FuzzingCornelius Aschermann, Sergej Schumilo, Ali Abbasi, Thorsten HolzS&P 2020 · 被引用 146 次
- GRIMOIRE: Synthesizing Structure while FuzzingTim Blazytko, Cornelius Aschermann, Moritz Schlögel, Ali Abbasi 等USENIX Security 2019 · 被引用 123 次
