Automated Attack Synthesis by Extracting Finite State Machines from Protocol Specification Documents
Maria Leonor Pacheco, Max von Hippel, Ben Weintraub, Dan Goldwasser, Cristina Nita-Rotaru
Abstract
Automated attack discovery techniques, such as attacker synthesis or model-based fuzzing, provide powerful ways to ensure network protocols operate correctly and securely. Such techniques, in general, require a formal representation of the protocol, often in the form of a finite state machine (FSM). Unfortunately, many protocols are only described in English prose, and implementing even a simple network protocol as an FSM is time-consuming and prone to subtle logical errors. Automatically extracting protocol FSMs from documentation can significantly contribute to increased use of these techniques and result in more robust and secure protocol implementations.In this work we focus on attacker synthesis as a representative technique for protocol security, and on RFCs as a representative format for protocol prose description. Unlike other works that rely on rule-based approaches or use off-the-shelf NLP tools directly, we suggest a data-driven approach for extracting FSMs from RFC documents. Specifically, we use a hybrid approach consisting of three key steps: (1) large-scale word-representation learning for technical language, (2) focused zero-shot learning for mapping protocol text to a protocol-independent information language, and (3) rule-based mapping from protocol-independent information to a specific protocol FSM. We show the generalizability of our FSM extraction by using the RFCs for six different protocols: BGPv4, DCCP, LTP, PPTP, SCTP and TCP. We demonstrate how automated extraction of an FSM from an RFC can be applied to the synthesis of attacks, with TCP and DCCP as case-studies. Our approach shows that it is possible to automate attacker synthesis against protocols by using textual specifications such as RFCs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers16
- Hermes: Unlocking Security Analysis of Cellular Network Protocols by Synthesizing Finite State Machines from Natural Language SpecificationsAbdullah Al Ishtiaq, Sarkar Snigdha Sarathi Das, Syed Md. Mukit Rashid, Ali Ranjbar et al.USENIX Security 2024 · 28 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- CellularLint: A Systematic Approach to Identify Inconsistent Behavior in Cellular Network SpecificationsMirza Masfiqur Rahman, Imtiaz Karim, Elisa BertinoUSENIX Security 2024 · 17 citations
- Formal Model-Driven Analysis of Resilience of GossipSub to Attacks from Misbehaving PeersAnkit Kumar, Max von Hippel, Panagiotis Manolios, Cristina Nita-RotaruS&P 2024 · 8 citations
- FORAY: Towards Effective Attack Synthesis against Deep Logical Vulnerabilities in DeFi ProtocolsHongbo Wen, Hanzhi Liu, Jiaxin Song, Yanju Chen et al.CCS 2024 · 6 citations
Builds on9
- SmartAuth: User-Centered Authorization for the Internet of ThingsYuan Tian, Nan Zhang, Yue-Hsun Lin, XiaoFeng Wang et al.USENIX Security 2017 · 231 citations
- On the Safety of IoT Device Physical Interaction ControlWenbo Ding, Hongxin HuCCS 2018 · 169 citations
- Towards the Detection of Inconsistencies in Public Security Vulnerability ReportsYing Dong, Wenbo Guo, Yueqi Chen, Xinyu Xing et al.USENIX Security 2019 · 149 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Automated Attack Discovery in TCP Congestion Control Using a Model-guided ApproachSamuel Jero, Md. Endadul Hoque, David R. Choffnes, Alan Mislove et al.NDSS 2018 · 46 citations
Related papers
- Generating Precise Format Specification for Network Protocols Through Adversarial LLM InteractionsHengdi Ye, Bing Shui, Jielun Wu, Yufan Zhou et al.USENIX Security 2026
- Large Language Model guided Protocol FuzzingRuijie Meng, Martin Mirchev, Marcel Böhme, Abhik RoychoudhuryNDSS 2024
- LLMs Unleashed: Generating Protocol Code from RFC SpecificationsJunfeng Long, Jinshu Su, Biao HanAAAI 2026
- Automated Construction of High-Quality Initial Seed Corpus for Network Protocol FuzzingWeicheng Lin, Laile Xi, Yaowen Zheng, Shenghao Lin et al.INFOCOM 2026
- SemFuzz: A Semantics-Aware Fuzzing Framework for Network Protocol ImplementationsYanbang Sun, Quan Luo, Yuelin Wang, Qian Chen et al.WWW 2026
