Large Language Model guided Protocol Fuzzing
Ruijie Meng, Martin Mirchev, Marcel Böhme, Abhik Roychoudhury
摘要
—How to find security flaws in a protocol implementation without a machine-readable specification of the protocol? Facing the internet, protocol implementations are particularly security-critical software systems where inputs must adhere to a specific structure and order that is often informally specified in hundreds of pages in natural language (RFC). Without some machine-readable version of that protocol, it is difficult to automatically generate valid test inputs for its implementation that follow the required structure and order. It is possible to partially alleviate this challenge using mutational fuzzing on a set of recorded message sequences as seed inputs. However, the set of available seeds is often quite limited and will hardly cover the great diversity of protocol states and input structures. In this paper, we explore the opportunities of systematic interaction with pre-trained large language models (LLMs), which have ingested millions of pages of human-readable protocol specifications, to draw out machine-readable information about the protocol that can be used during protocol fuzzing. We use the knowledge of the LLMs about protocol message types for well-known protocols. We also checked the LLM’s capability in detecting “states” for stateful protocol implementations by generating sequences of messages and predicting response codes. Based on these observations, we have developed an LLM-guided protocol implementation fuzzing engine. Our protocol fuzzer C HAT AFL constructs grammars for each message type in a protocol, and then mutates messages or predicts the next messages in a message sequence via interactions with LLMs. Experiments on a wide range of real-world protocols from P RO F UZZ B ENCH show significant efficacy in state and code coverage. Our LLM-guided stateful fuzzer was compared with state-of-the-art fuzzers AFLN ET and NSF UZZ . C HAT AFL covers 47.60% and 42.69% more state transitions, 29.55% and 25.75% more
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper66
- SelfCodeAlign: Self-Alignment for Code GenerationYuxiang Wei, Federico Cassano, Jiawei Liu, Yifeng Ding 等NeurIPS 2024 · 被引用 79 次
- WhiteFox: White-Box Compiler Fuzzing Empowered by Large Language ModelsChenyuan Yang, Yinlin Deng, Runyu Lu, Jiayi Yao 等OOPSLA 2024 · 被引用 74 次
- LLMIF: Augmented Large Language Model for Fuzzing IoT DevicesJincheng Wang, Le Yu, Xiapu LuoS&P 2024 · 被引用 61 次
- Fuzzing BusyBox: Leveraging LLM and Crash Reuse for Embedded Bug UnearthingAsmita, Yaroslav Oliinyk, Michael Scott, Ryan Tsang 等USENIX Security 2024 · 被引用 56 次
- From One Thousand Pages of Specification to Unveiling Hidden Bugs: Large Language Model Assisted Fuzzing of Matter IoT DevicesXiaoyue Ma, Lannan Luo, Qiang ZengUSENIX Security 2024 · 被引用 49 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Directed Greybox FuzzingMarcel Böhme, Van-Thuan Pham, Manh-Dung Nguyen, Abhik RoychoudhuryCCS 2017 · 被引用 836 次
- Evaluating Fuzz TestingGeorge Klees, Andrew Ruef, Benji Cooper, Shiyi Wei 等CCS 2018 · 被引用 753 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
相关 Paper
- Generating Precise Format Specification for Network Protocols Through Adversarial LLM InteractionsHengdi Ye, Bing Shui, Jielun Wu, Yufan Zhou 等USENIX Security 2026
- An LLM-Guided Fuzzing of Proprietary Industrial Communication Protocols with Context KnowledgeTianci Pan, Huan Qian, Yaowen Zheng, Haining Wang 等INFOCOM 2026
- Automated Construction of High-Quality Initial Seed Corpus for Network Protocol FuzzingWeicheng Lin, Laile Xi, Yaowen Zheng, Shenghao Lin 等INFOCOM 2026
- Low-Cost and Comprehensive Non-textual Input Fuzzing with LLM-Synthesized Input GeneratorsKunpeng Zhang, Zongjie Li, Daoyuan Wu, Shuai Wang 等USENIX Security 2025
- SemFuzz: A Semantics-Aware Fuzzing Framework for Network Protocol ImplementationsYanbang Sun, Quan Luo, Yuelin Wang, Qian Chen 等WWW 2026
