Thinking More, Harnessing Better: Automatic Harness Generation with Dataflow Aggregation and Workflow Decomposition
Xing Zhang, Zikang Huang, Gang Yang, CongChong Wang, Lu Liu, Bin Yin, Mingyi Wang, Ziquan Zhao, Min Li, Zhenyu Chen, Bo Wu, Lingyun Ying
摘要
High-quality fuzz harnesses are essential for effective gray-box fuzzing. While Large Language Models (LLMs) offer promise for automating this task, existing one-turn generation methods suffer from hallucinations and inadequate coverage due to coarsegrained function targeting and misaligned generation workflows. We present SynapseFlow, an automatic harness generator that addresses these limitations through two key innovations: dataflowaware function aggregation and a staged, rollback-enabled generation workflow decomposition. SynapseFlow first analyzes source code to construct Structural Flow Graphs and extract coherent Function Triplets. It then synthesizes harnesses via a decomposed fourstage process governed by a staged rollback algorithm to ensure correctness. We evaluated SynapseFlow on 25 real-world opensource software projects. The experimental results indicate that SynapseFlow outperforms state-of-the-art tools (OSS-Fuzz-Gen, CKGFuzzer, PromeFuzz), achieving 3.07×, 1.71×, and 4.26× higher branch coverage, and 1.77×, 1.51×, and 1.36× higher bug detection rates, respectively. Most importantly, SynapseFlow discovered 7 previously unreported bugs (5 assigned CVEs), demonstrating its practical effectiveness in real-world bug discovery. CCS Concepts • Security and privacy → Software security engineering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Fuzz4All: Universal Fuzzing with Large Language ModelsChunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel 等ICSE 2024 · 被引用 155 次
- Large Language Models are Edge-Case Generators: Crafting Unusual Programs for Fuzzing Deep Learning LibrariesYinlin Deng, Chunqiu Steven Xia, Chenyuan Yang, Shizhuo Dylan Zhang 等ICSE 2024 · 被引用 85 次
- WhiteFox: White-Box Compiler Fuzzing Empowered by Large Language ModelsChenyuan Yang, Yinlin Deng, Runyu Lu, Jiayi Yao 等OOPSLA 2024 · 被引用 74 次
- SoK: Prudent Evaluation Practices for FuzzingMoritz Schloegel, Nils Bars, Nico Schiller, Lukas Bernhard 等S&P 2024 · 被引用 69 次
- APICraft: Fuzz Driver Generation for Closed-source SDK LibrariesCen Zhang, Xingwei Lin, Yuekang Li, Yinxing Xue 等USENIX Security 2021 · 被引用 64 次
相关 Paper
- PromeFuzz: A Knowledge-Driven Approach to Fuzzing Harness Generation with Large Language ModelsYuwei Liu, Junquan Deng, Xiangkun Jia, Yanhao Wang 等CCS 2025
- ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer SpaceChuyang Chen, Brendan Dolan-Gavitt, Zhiqiang LinUSENIX Security 2025
- WildSync: Automated Fuzzing Harness Synthesis via Wild API Usage RecoveryWei-Cheng Wu, Stefan Nagy, Christophe HauserISSTA 2025 · 被引用 1 次
- Automatic, Expressive, and Scalable Fuzzing with StitchingHarrison Green, Fraser Brown, Claire Le GouesCCS 2026 · 被引用 1 次
- Thought Is All You Need: Smart Contract Vulnerability Detection with Thought-Augmented Large Language ModelChaoyuan Peng, Muhui Jiang, Yajin Zhou, Lei WuFSE 2026
