3DGen: AI-Assisted Generation of Provably Correct Binary Format Parsers
Sarah Fakhoury, Markus Kuppe, Shuvendu K. Lahiri, Tahina Ramananandro, Nikhil Swamy
摘要
Improper parsing of attacker-controlled input is a leading source of software security vulnerabilities, especially when programmers transcribe informal format descriptions in RFCs into efficient parsing logic in low-level, memory unsafe languages. Several researchers have proposed formal specification languages for data formats from which efficient code can be extracted. However, distilling informal requirements into formal specifications is challenging and, despite their benefits, new, formal languages are hard for people to learn and use. In this work, we present 3DGen, a framework that makes use of AI agents to transform mixed informal input, including natural language documents (i.e., RFCs) and example inputs into format specifications in a language called 3D. To support humans in understanding and trusting the generated specifications, 3DGen uses symbolic methods to also synthesize test inputs that can be validated against an external oracle. Symbolic test generation also helps in distinguishing multiple plausible solutions. Through a process of repeated refinement, 3DGen produces a 3D specification that conforms to a test suite, and which yields safe, efficient, provably correct, parsing code in C. We have evaluated 3DGen on 20 Internet standard formats, demonstrating the potential for AI-agents to produce formally verified <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> code at a non-trivial scale. A key enabler is the use of a domain-specific language to limit AI outputs to a class for which automated. symbolic analysis is tractable.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- AutoVerus: Automated Proof Generation for Rust CodeChenyuan Yang, Xuheng Li, Md Rakib Hossain Misu, Jianan Yao 等OOPSLA 2025 · 被引用 11 次
- Towards Functional Correctness of Large Code Models with Selective GenerationJaewoo Jeong, Taesoo Kim, Sangdon ParkICML 2026
- Generating Precise Format Specification for Network Protocols Through Adversarial LLM InteractionsHengdi Ye, Bing Shui, Jielun Wu, Yufan Zhou 等USENIX Security 2026
它引用的顶会 Paper8
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code ContributionsHammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt 等S&P 2022 · 被引用 725 次
- EverParse: Verified Secure Zero-Copy Parsers for Authenticated Message FormatsTahina Ramananandro, Antoine Delignat-Lavaud, Cédric Fournet, Nikhil Swamy 等USENIX Security 2019 · 被引用 70 次
- CodeT: Code Generation with Generated TestsBei Chen, Fengji Zhang, Anh Nguyen, Daoguang Zan 等ICLR 2023 · 被引用 64 次
- Semi-automated protocol disambiguation and code generationJane Yen, Tamás Lévai, Qinyuan Ye, Xiang Ren 等SIGCOMM 2021 · 被引用 33 次
相关 Paper
- FormalJudge: A Neuro-Symbolic Paradigm for Agentic OversightJiayi Zhou, Yang Sheng, Hantao Lou, Yaodong Yang 等ICML 2026
- SpecGen: Automated Generation of Formal Program Specifications via Large Language ModelsLezhi Ma, Shangqing Liu, Yi Li, Xiaofei Xie 等ICSE 2025 · 被引用 25 次
- Hardening attack surfaces with formally proven binary format parsersNikhil Swamy, Tahina Ramananandro, Aseem Rastogi, Irina Spiridonova 等PLDI 2022 · 被引用 18 次
- LLMs Unleashed: Generating Protocol Code from RFC SpecificationsJunfeng Long, Jinshu Su, Biao HanAAAI 2026
- PAT-Agent: Autoformalization for Model CheckingXinyue Zuo, Yifan Zhang, Hongshu Wang, Yufan Cai 等ASE 2025 · 被引用 1 次
