Specializing Language Models for Textual Fuzzing via Reinforcement Learning
Jiayi Lin, Liangcai Su, Junzhe Li, Chenxiong Qian
Abstract
Fuzzing is effective for vulnerability discovery but struggles with complex targets such as compilers, interpreters, and database engines, which accept textual input that must satisfy intricate syntactic and semantic constraints. Although language models (LMs) have attracted interest for this task due to their vast latent knowledge and reasoning potential, their practical adoption has been limited. The major challenges stem from insufficient exploration of deep program logic among real-world codebases, and the high cost of leveraging larger models. To overcome these challenges, we propose R1-Fuzz, the first framework that leverages reinforcement learning (RL) to train cost-efficient LMs and integrate them for complex textual fuzzing. R1-FUZZ introduces two key designs: coverage-slicing-based question construction and a distance-based reward calculation. Using these methods, R1-Fuzz performs RL-based post-training on a constructed dataset of questions, utilizing our fine-grained reward signals. We use R1-FUZZ to train a cost-efficient LM named R1-Fuzz-7B, derived from the Qwen2.5-7B model, to excel at fuzzing input generation. Then, R1-Fuzz designs a fuzzing workflow that tightly integrates LMs to reason deep program semantics during practical fuzzing. Evaluations on diverse real-world targets show that R1-FUZZ-7B can rival or even outperform much larger models (e.g., GPT-o4-mini, Deepseek-V3) in real-world fuzzing, achieving up to 75% higher coverage than state-of-the-art fuzzers and discovering 37 previously unknown vulnerabilities, demonstrating its practicality.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 55aed2a7-77ad-408b-b3bf-420250f4cbc1Related papers
- Fuzzing JavaScript Interpreters with Coverage-Guided Reinforcement Learning for LLM-Based MutationJueon Eom, Seyeon Jeong, Taekyoung KwonISSTA 2024 · 26 citations
- GenHuzz: An Efficient Generative Hardware FuzzerLichao Wu, Mohamadreza Rostami, Huimin Li, Jeyavijayan Rajendran et al.USENIX Security 2025
- ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer SpaceChuyang Chen, Brendan Dolan-Gavitt, Zhiqiang LinUSENIX Security 2025
- Low-Cost and Comprehensive Non-textual Input Fuzzing with LLM-Synthesized Input GeneratorsKunpeng Zhang, Zongjie Li, Daoyuan Wu, Shuai Wang et al.USENIX Security 2025
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel et al.ICSE 2024 · 155 citations
