Expecto: Extracting Formal Specifications from Natural Language Description for Trustworthy Oracles
Dongjae Lee, Kihong Heo
摘要
Specification-Driven Development (SDD) has emerged as a promising paradigm in software development. This trend is fueled by recent advances in leveraging large language models (LLMs) to generate code from user intents expressed in natural language. However, the reliance on natural language specifications introduces ambiguity and challenges in ensuring correctness. To address these problems, we present Expecto, a system that automatically extracts trustworthy formal specifications from natural language intents. Expecto employs a neuro-symbolic approach, combining the strengths of LLMs and program synthesis techniques. Our key idea is to adopt a top-down, modular specification synthesis algorithm that breaks down the complex task of specification extraction into manageable units. The specifications are written in a domain-specific language (DSL) designed to succinctly express formal specifications of functional requirements. This modular synthesis with a concise DSL reduces the reasoning complexity for LLMs and enhances the accuracy of the extracted specifications. Our results demonstrate that Expecto significantly improves the accuracy and reliability of extracted specifications compared to a monolithic, purely LLM-based approach. Furthermore, when applied to real-world buggy programs in Defects4J, Expecto successfully generates formal specifications that detect more bugs than the baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
- Autoformalization with Large Language ModelsYuhuai Wu, Albert Qiaochu Jiang, Wenda Li, Markus N. Rabe 等NeurIPS 2022 · 被引用 364 次
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei 等ICLR 2023 · 被引用 318 次
- Don't Trust: Verify - Grounding LLM Quantitative Reasoning with AutoformalizationJin Peng Zhou, Charles Staats, Wenda Li, Christian Szegedy 等ICLR 2024 · 被引用 72 次
相关 Paper
- Can Large Language Models Transform Natural Language Intent into Formal Method Postconditions?Madeline Endres, Sarah Fakhoury, Saikat Chakraborty, Shuvendu K. LahiriFSE 2024 · 被引用 27 次
- FormalJudge: A Neuro-Symbolic Paradigm for Agentic OversightJiayi Zhou, Yang Sheng, Hantao Lou, Yaodong Yang 等ICML 2026
- SpecGen: Automated Generation of Formal Program Specifications via Large Language ModelsLezhi Ma, Shangqing Liu, Yi Li, Xiaofei Xie 等ICSE 2025 · 被引用 25 次
- SpecMind: Cognitively Inspired, Interactive Multi-Turn Framework for Postcondition InferenceCuong Chi Le, Minh V. T. Pham, Tung Duy Vu, Cuong Duc Van 等ACL 2026 · 被引用 2 次
- Bridging Natural Language and Formal Specification-Automated Translation of Software Requirements to LTL via Hierarchical Semantics Decomposition Using LLMsZhi Ma, Cheng Wen, Zhexin Su, Xiao Liang 等ASE 2025 · 被引用 3 次
