Automatic, Expressive, and Scalable Fuzzing with Stitching
Harrison Green, Fraser Brown, Claire Le Goues
摘要
Fuzzing is a powerful technique for finding bugs in software libraries, but scaling it remains difficult. Automated harness generation commits to fixed API sequences at synthesis time, limiting the behaviors each harness can test. Approaches that instead explore new sequences dynamically lack the expressiveness to model real-world usage constraints leading to false positives from straightforward API misuse.
We propose stitching, a technique that encodes API usage constraints in pieces that a fuzzer dynamically assembles at runtime. A static type system governs how objects flow between blocks, while a dynamically-checked extrinsic typestate tracks arbitrary metadata across blocks, enabling specifications to express rich semantic constraints such as object state dependencies and cross-function preconditions. This allows a single specification to describe an open-ended space of valid API interactions that the fuzzer explores guided by coverage feedback.
We implement stitching in STITCH, using LLMs to automatically configure projects for fuzzing, synthesize a specification, triage crashes, and repair the specification itself. We evaluated STITCH against four state-of-the-art tools on 33 benchmarks, where it achieved the highest code coverage on 21 and found 30 true-positive bugs compared to 10 by all other tools combined, with substantially higher precision (70% vs. 12% for the next-best LLM-based tool). Deployed automatically on 1365 widely used open-source projects, STITCH discovered 131 new bugs across 102 projects, 73 of which have already been patched.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LibAFL: A Framework to Build Modular and Reusable FuzzersAndrea Fioraldi, Dominik Christian Maier, Dongjia Zhang, Davide BalzarottiCCS 2022 · 被引用 71 次
- GraphFuzz: Library API Fuzzing with Lifetime-aware Dataflow GraphsHarrison Green, Thanassis AvgerinosICSE 2022 · 被引用 38 次
- CAT-LM Training Language Models on Aligned Code And TestsNikitha Rao, Kush Jain, Uri Alon, Claire Le Goues 等ASE 2023 · 被引用 37 次
- Prompt Fuzzing for Fuzz Driver GenerationYunlong Lyu, Yuxuan Xie, Peng Chen, Hao ChenCCS 2024 · 被引用 21 次
相关 Paper
- WildSync: Automated Fuzzing Harness Synthesis via Wild API Usage RecoveryWei-Cheng Wu, Stefan Nagy, Christophe HauserISSTA 2025 · 被引用 1 次
- No Harness, No Problem: Oracle-guided Harnessing for Auto-generating C API Fuzzing HarnessesGabriel Sherman, Stefan NagyICSE 2025 · 被引用 1 次
- PromeFuzz: A Knowledge-Driven Approach to Fuzzing Harness Generation with Large Language ModelsYuwei Liu, Junquan Deng, Xiangkun Jia, Yanhao Wang 等CCS 2025
- FuzzGen: Automatic Fuzzer GenerationKyriakos K. Ispoglou, Daniel Austin, Vishwath Mohan, Mathias PayerUSENIX Security 2020
- Thinking More, Harnessing Better: Automatic Harness Generation with Dataflow Aggregation and Workflow DecompositionXing Zhang, Zikang Huang, Gang Yang, CongChong Wang 等CCS 2026
