Can Large Language Models Transform Natural Language Intent into Formal Method Postconditions?
Madeline Endres, Sarah Fakhoury, Saikat Chakraborty, Shuvendu K. Lahiri
Abstract
Informal natural language that describes code functionality, such as code comments or function documentation, may contain substantial information about a program’s intent. However, there is typically no guarantee that a program’s implementation and natural language documentation are aligned. In the case of a conflict, leveraging information in code-adjacent natural language has the potential to enhance fault localization, debugging, and code trustworthiness. In practice, however, this information is often underutilized due to the inherent ambiguity of natural language, which makes natural language intent challenging to check programmatically. The “emergent abilities” of Large Language Models (LLMs) have the potential to facilitate the translation of natural language intent to programmatically checkable assertions. However, it is unclear if LLMs can correctly translate informal natural language specifications into formal specifications that match programmer intent. Additionally, it is unclear if such translation could be useful in practice. In this paper, we describe nl2postcond , the problem of leveraging LLMs for transforming informal natural language to formal method postconditions, expressed as program assertions. We introduce and validate metrics to measure and compare different nl2postcond approaches, using the correctness and discriminative power of generated postconditions. We then use qualitative and quantitative methods to assess the quality of nl2postcond postconditions, finding that they are generally correct and able to discriminate incorrect code. Finally, we find that nl2postcond via LLMs has the potential to be helpful in practice; nl2postcond generated postconditions were able to catch 64 real-world historical bugs from Defects4J .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext acb17321-274e-4be5-896c-cf9dff7bb489Cited by top-tier papers22
- VERINA: Benchmarking Verifiable Code GenerationZhe Ye, Zhengxu Yan, Jingxuan He, Timothe Kasriel et al.ICLR 2026 · 34 citations
- SpecRover: Code Intent Extraction via LLMsHaifeng Ruan, Yuntong Zhang, Abhik RoychoudhuryICSE 2025 · 12 citations
- COFFE: A Code Efficiency Benchmark for Code GenerationYun Peng, Jun Wan, Yichen Li, Xiaoxue RenFSE 2025 · 8 citations
- Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code GenerationHaoxiang Jia, Robbie Morris, He Ye, Federica Sarro et al.ASE 2025 · 6 citations
- PGS: Effective LLM Code Refinement via Property-Oriented and Structurally Minimal FeedbackLehan He, Zeren Chen, Zhe Zhang, Xiang Gao et al.ICML 2026 · 4 citations
Builds on11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program SynthesisErik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu et al.ICLR 2023 · 234 citations
- CodaMosa: Escaping Coverage Plateaus in Test Generation with Pre-trained Large Language ModelsCaroline Lemieux, Jeevana Priya Inala, Shuvendu K. Lahiri, Siddhartha SenICSE 2023 · 221 citations
Related papers
- Expecto: Extracting Formal Specifications from Natural Language Description for Trustworthy OraclesDongjae Lee, Kihong HeoPLDI 2026
- SpecMind: Cognitively Inspired, Interactive Multi-Turn Framework for Postcondition InferenceCuong Chi Le, Minh V. T. Pham, Tung Duy Vu, Cuong Duc Van et al.ACL 2026 · 2 citations
- Measuring the Influence of Incorrect Code on Test GenerationDong Huang, Jie M. Zhang, Mark Harman, Mingzhe Du et al.ICSE 2026
- Large Language Models are Few-Shot Summarizers: Multi-Intent Comment Generation via In-Context LearningMingyang Geng, Shangwen Wang, Dezun Dong, Haotian Wang et al.ICSE 2024 · 124 citations
- HoarePrompt: Structural Reasoning About Program Correctness in Natural LanguageDimitrios Stamatios Bouras, Yihan Dai, Tairan Wang, Yingfei Xiong et al.ICSE 2026
