SnakeCharmer: Automatic Fuzzing Harness Generation for Pure and Hybrid Python Libraries
Gabriel Sherman, Stefan Nagy
摘要
With Python's rising popularity, ensuring the correctness of its ever-growing ecosystem of software libraries is more critical than ever. Recently, fuzzing has become a de facto technique for vetting software libraries, enabled via the use of harnesses: small wrapper programs that inject fuzzer-generated test cases into the library under test. While harness creation has shed its reliance on human expertise and is now fully automated for languages such as C and C++, Python remains uniquely challenging-both for pure Python libraries as well as hybrid ones combining Python with native C/C++ extensions-due to (1) limited visibility across language boundaries, (2) the absence of reliable bug oracles, and (3) incomplete type information. Consequently, attempts at automating harnessing for Python fail to both uphold critical runtime behaviors and produce the structured call and data flows needed for effective fuzzing, leaving much of today's Python ecosystem largely unvetted.
To overcome these challenges and broaden fuzzing's reach across Python libraries, this paper introduces SnakeCharmer: the first automated harness generation approach for both pure and hybrid Python libraries. At its core, SnakeCharmer leverages static analysis to first capture key interface information from both Python and native code components, subsequently enriching it with runtime-captured type information and exception behaviors. During fuzzing, SnakeCharmer further distinguishes between expected exceptions and true library bugs, filtering out benign exceptions that would otherwise derail testing progress. Together, these techniques significantly enhance the scope and effectiveness of fuzzing across the Python library ecosystem, enabling the automated discovery of bugs in code previously inaccessible to existing Python fuzzing efforts.
We evaluate SnakeCharmer alongside today's leading Python auto-harnessing approach, PyRTFuzz; the actively fuzzed expert-written harnesses from both OSS-Fuzz and PolyFuzz; and the harnesses generated by Google's own state-of-the-art LLM-driven automatic harnessing approach, OSS-Fuzz-Gen. Across 21 diverse Python libraries, SnakeCharmer attains type-recovery precision and exception-filtering accuracy of 95% and 97%, respectively, further attaining 1.48×, 1.87×, 1.78×, and 1.40× the code coverage of the fuzzing harnesses from PyRTFuzz, OSS-Fuzz, PolyFuzz, and OSS-Fuzz-Gen, respectively. Further, SnakeCharmer finds 16, 24, and 24 more Python library bugs than all expert-and LLM-created harnesses as well as PyRTFuzz, respectively-uncovering a total of 20 new bugs, with 18 since confirmed or fixed by developers.
CCS Concepts: • Software and its engineering → Software testing and debugging.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei 等CCS 2018 · 被引用 753 次
- Seed selection for successful fuzzingAdrian Herrera, Hendra Gunadi, Shane Magrath, Michael Norrish 等ISSTA 2021 · 被引用 95 次
- APICraft: Fuzz Driver Generation for Closed-source SDK LibrariesCen Zhang, Xingwei Lin, Yuekang Li, Yinxing Xue 等USENIX Security 2021 · 被引用 64 次
- GraphFuzz: Library API Fuzzing with Lifetime-aware Dataflow GraphsHarrison Green, Thanassis AvgerinosICSE 2022 · 被引用 38 次
- The evolution of type annotations in python: an empirical studyLuca Di Grazia, Michael PradelFSE 2022 · 被引用 29 次
相关 Paper
- No Harness, No Problem: Oracle-guided Harnessing for Auto-generating C API Fuzzing HarnessesGabriel Sherman, Stefan NagyICSE 2025 · 被引用 1 次
- WildSync: Automated Fuzzing Harness Synthesis via Wild API Usage RecoveryWei-Cheng Wu, Stefan Nagy, Christophe HauserISSTA 2025 · 被引用 1 次
- PyRTFuzz: Detecting Bugs in Python Runtimes via Two-Level Collaborative FuzzingWen Li, Haoran Yang, Xiapu Luo, Long Cheng 等CCS 2023 · 被引用 14 次
- Liberating Libraries through Automated Fuzz Driver Generation: Striking a Balance without Consumer CodeFlavio Toffalini, Nicolas Badoux, Zurab Tsinadze, Mathias PayerFSE 2025
- Automatic, Expressive, and Scalable Fuzzing with StitchingHarrison Green, Fraser Brown, Claire Le GouesCCS 2026 · 被引用 1 次
