Validating Formal Specifications with LLM-Generated Test Cases
Alcino Cunha, Nuno Macedo
摘要
Abstract Validation is a central activity when developing formal specifications. Similarly to coding, a possible validation technique is to define upfront test cases or scenarios that a future specification should satisfy or not. Unfortunately, specifying such test cases is burdensome and error prone, which could cause users to skip this validation task. This paper reports the results of an empirical evaluation of using pre-trained large language models (LLMs) to automate the generation of test cases from natural language requirements. In particular, we focus on test cases for structural requirements of simple domain models formalized in the Alloy specification language. Our evaluation focuses on the state-of-the-art GPT-5 model, but results from other closed- and open-source LLMs are also reported. The results show that, in this context, GPT-5 is already quite effective at generating positive (and negative) test cases that are syntactically correct and that satisfy (or not) the given requirement, and that can detect many wrong specifications written by humans.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- On the Evaluation of Large Language Models in Unit Test GenerationLin Yang, Chen Yang, Shutao Gao, Weijing Wang 等ASE 2024 · 被引用 42 次
- LLM4Fin: Fully Automating LLM-Powered Test Case Generation for FinTech Software Acceptance TestingZhiyi Xue, Liangguo Li, Senyue Tian, Xiaohong Chen 等ISSTA 2024 · 被引用 19 次
- PAT-Agent: Autoformalization for Model CheckingXinyue Zuo, Yifan Zhang, Hongshu Wang, Yufan Cai 等ASE 2025 · 被引用 1 次
相关 Paper
- Guiding Enumerative Program Synthesis with Large Language ModelsYixuan Li, Julian Parsert, Elizabeth PolgreenCAV 2024 · 被引用 19 次
- Evaluating and Improving ChatGPT for Unit Test GenerationZhiqiang Yuan, Mingwei Liu, Shiji Ding, Kaixin Wang 等FSE 2024 · 被引用 89 次
- A Tale of 1001 LoC: Potential Runtime Error-Guided Specification Synthesis for Verifying Large-Scale ProgramsZhongyi Wang, Tengjie Lin, Mingshuai Chen, Haokun Li 等OOPSLA 2026 · 被引用 1 次
- SpecGen: Automated Generation of Formal Program Specifications via Large Language ModelsLezhi Ma, Shangqing Liu, Yi Li, Xiaofei Xie 等ICSE 2025 · 被引用 25 次
- Fill in the Blank: Context-aware Automated Text Input Generation for Mobile GUI TestingZhe Liu, Chunyang Chen, Junjie Wang, Xing Che 等ICSE 2023 · 被引用 107 次
