3DGen: AI-Assisted Generation of Provably Correct Binary Format Parsers
Sarah Fakhoury, Markus Kuppe, Shuvendu K. Lahiri, Tahina Ramananandro, Nikhil Swamy
Abstract
Improper parsing of attacker-controlled input is a leading source of software security vulnerabilities, especially when programmers transcribe informal format descriptions in RFCs into efficient parsing logic in low-level, memory unsafe languages. Several researchers have proposed formal specification languages for data formats from which efficient code can be extracted. However, distilling informal requirements into formal specifications is challenging and, despite their benefits, new, formal languages are hard for people to learn and use. In this work, we present 3DGen, a framework that makes use of AI agents to transform mixed informal input, including natural language documents (i.e., RFCs) and example inputs into format specifications in a language called 3D. To support humans in understanding and trusting the generated specifications, 3DGen uses symbolic methods to also synthesize test inputs that can be validated against an external oracle. Symbolic test generation also helps in distinguishing multiple plausible solutions. Through a process of repeated refinement, 3DGen produces a 3D specification that conforms to a test suite, and which yields safe, efficient, provably correct, parsing code in C. We have evaluated 3DGen on 20 Internet standard formats, demonstrating the potential for AI-agents to produce formally verified <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> code at a non-trivial scale. A key enabler is the use of a domain-specific language to limit AI outputs to a class for which automated. symbolic analysis is tractable.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dbe93856-f165-48de-a43f-d4b9cc9e8d1dCited by top-tier papers3
- AutoVerus: Automated Proof Generation for Rust CodeChenyuan Yang, Xuheng Li, Md Rakib Hossain Misu, Jianan Yao et al.OOPSLA 2025 · 11 citations
- Towards Functional Correctness of Large Code Models with Selective GenerationJaewoo Jeong, Taesoo Kim, Sangdon ParkICML 2026
- Generating Precise Format Specification for Network Protocols Through Adversarial LLM InteractionsHengdi Ye, Bing Shui, Jielun Wu, Yufan Zhou et al.USENIX Security 2026
Builds on8
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt et al.S&P 2022 · 725 citations
- EverParse: Verified Secure Zero-Copy Parsers for Authenticated Message FormatsTahina Ramananandro, Antoine Delignat-Lavaud, Cédric Fournet, Nikhil Swamy et al.USENIX Security 2019 · 70 citations
- CodeT: Code Generation with Generated TestsBei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan et al.ICLR 2023 · 64 citations
- Semi-automated protocol disambiguation and code generationJane Yen, Tamás Lévai, Qinyuan Ye, Xiang Ren et al.SIGCOMM 2021 · 33 citations
Related papers
- FormalJudge: A Neuro-Symbolic Paradigm for Agentic OversightJiayi Zhou, Yang Sheng, Hantao Lou, Yaodong Yang et al.ICML 2026
- SpecGen: Automated Generation of Formal Program Specifications via Large Language ModelsLezhi Ma, Shangqing Liu, Yi Li, Xiaofei Xie et al.ICSE 2025 · 25 citations
- Hardening attack surfaces with formally proven binary format parsersNikhil Swamy, Tahina Ramananandro, Aseem Rastogi, Irina Spiridonova et al.PLDI 2022 · 18 citations
- LLMs Unleashed: Generating Protocol Code from RFC SpecificationsJunfeng Long, Jinshu Su, Biao HanAAAI 2026
- PAT-Agent: Autoformalization for Model CheckingXinyue Zuo, Yifan Zhang, Hongshu Wang, Yufan Cai et al.ASE 2025 · 1 citation
