Lune

S&P2026顶会

Specializing Language Models for Textual Fuzzing via Reinforcement Learning

Jiayi Lin, Liangcai Su, Junzhe Li, Chenxiong Qian

2026年份

摘要

Fuzzing is effective for vulnerability discovery but struggles with complex targets such as compilers, interpreters, and database engines, which accept textual input that must satisfy intricate syntactic and semantic constraints. Although language models (LMs) have attracted interest for this task due to their vast latent knowledge and reasoning potential, their practical adoption has been limited. The major challenges stem from insufficient exploration of deep program logic among real-world codebases, and the high cost of leveraging larger models. To overcome these challenges, we propose R1-Fuzz, the first framework that leverages reinforcement learning (RL) to train cost-efficient LMs and integrate them for complex textual fuzzing. R1-FUZZ introduces two key designs: coverage-slicing-based question construction and a distance-based reward calculation. Using these methods, R1-Fuzz performs RL-based post-training on a constructed dataset of questions, utilizing our fine-grained reward signals. We use R1-FUZZ to train a cost-efficient LM named R1-Fuzz-7B, derived from the Qwen2.5-7B model, to excel at fuzzing input generation. Then, R1-Fuzz designs a fuzzing workflow that tightly integrates LMs to reason deep program semantics during practical fuzzing. Evaluations on diverse real-world targets show that R1-FUZZ-7B can rival or even outperform much larger models (e.g., GPT-o4-mini, Deepseek-V3) in real-world fuzzing, achieving up to 75% higher coverage than state-of-the-art fuzzers and discovering 37 previously unknown vulnerabilities, demonstrating its practicality.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖