deepSURF: Detecting Memory Safety Vulnerabilities in Rust Through Fuzzing LLM-Augmented Harnesses
Georgios C. Androutsopoulos, Antonio Bianchi
Abstract
Although Rust ensures memory safety by default, it also permits the use of unsafe code, which can introduce memory safety vulnerabilities if misused. Unfortunately, existing tools for detecting memory bugs in Rust typically exhibit limited detection capabilities, inadequately handle Rust-specific types, or rely heavily on manual intervention. To address these limitations, we present deepSURF, a tool that integrates static analysis with Large Language Model (LLM)-guided fuzzing harness generation to effectively identify memory safety vulnerabilities in Rust libraries, specifically targeting unsafe code. deepSURF introduces a novel approach for handling generics by substituting them with custom types and generating tailored implementations for the required traits, enabling the fuzzer to simulate user-defined behaviors within the fuzzed library. Additionally, deepSURF employs LLMs to augment fuzzing harnesses dynamically, facilitating exploration of complex API interactions and significantly increasing the likelihood of exposing memory safety vulnerabilities. We evaluated deepSURF on 63 real-world Rust crates, successfully rediscovering 30 known memory safety bugs and uncovering 12 previously-unknown vulnerabilities (out of which have been assigned RustSec IDs and 3 have been patched), demonstrating clear improvements over state-of-the-art tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- RustGo: Fairly Directed Greybox Fuzzing for Enforcing Rust Memory SafetyDongyeon Yu, Jiun Min, Yewan Na, Mijung Kim et al.CCS 2026
- StepStone: LLM-Based GPU Kernel Driver Fuzzing via User-Space LibrariesXiaochen Zou, Juefei Pu, Arrdya Srivastav, Jonathan Cox et al.S&P 2026
Builds on19
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Large Language Models Are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language ModelsYinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang et al.ISSTA 2023 · 253 citations
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel et al.ICSE 2024 · 155 citations
- Understanding memory and thread safety practices and issues in real-world Rust programsBoqin Qin, Yilun Chen, Zeming Yu, Linhai Song et al.PLDI 2020 · 112 citations
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang et al.FSE 2024 · 89 citations
Related papers
- HarnessLLM: Rust Verification Harness Generation with Large Language ModelsMinghua Wang, Yuwei Liu, Lin HuangICSE 2026
- Crabtree: Rust API Test Synthesis Guided by Coverage and TypeYoshiki Takashima, Chanhee Cho, Ruben Martins, Limin Jia et al.OOPSLA 2024 · 3 citations
- CULPA: Universal Detection of Memory-Safety Bugs in Unsafe Rust Through the Lens of Safety RequirementsHung-Mao Chen, Bo Lu, Xu He, Xiaokuan Zhang et al.USENIX Security 2026
- RPG: Rust Library Fuzzing with Pool-based Fuzz Target Generation and Generic SupportZhiwu Xu, Bohao Wu, Cheng Wen, Bin Zhang et al.ICSE 2024 · 9 citations
- Unlocking a New Rust Programming Experience: Fast and Slow Thinking with LLMs to Conquer Undefined BehaviorsRenshuang Jiang, Pan Dong, Zhenling Duan, Yu Shi et al.DAC 2025
