Automating Requirements Formalization: Using LLMs and Low-Complexity Distinguishing Traces for Semantic Validation
Daniel Mendoza, Anastasia Mavridou, Andreas Katis, Caroline Trippel
Abstract
Translating natural language (NL) requirements into formal specifications is critical for verifying safety-critical systems, but it is error-prone and time-consuming when done manually. While Large Language Models (LLMs) can automate this translation, they often produce incorrect outputs that require extensive validation. In this paper, we propose ARTEMIS, an LLM-based framework that translates unstructured NL requirements into formal temporal logic (TL) specifications. Our framework reduces validation effort through three synergistic, automated techniques: (i) LLM Translation to Structured NL: We use LLMs to translate unstructured NL requirements into structured NL, which has an unambiguous mapping to TL. This intermediate representation reduces translation errors. (ii) Sub-Specification Generation: We generate low-complexity execution traces (i.e., system behaviors) that correspond to candidate specification fragments from the LLM translations. Users inspect these and accept or reject them. (iii) Balanced Distinguishing Trace Generation: We minimize the number of traces users need to inspect by pruning the candidate specification space. Each accepted or rejected trace eliminates candidates logarithmically. We evaluate ARTEMIS on five real-world safety-critical requirements datasets. The results show that it achieves 1.57X higher translation accuracy while reducing manual validation effort by up to 10.83X compared to state-of-the-art baselines.
• Software and its engineering → Formal methods; Requirements analysis; • Computing methodologies → Natural language processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1cb1c6e8-5fe3-465c-b83d-957663f524ceBuilds on7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng et al.ICLR 2024 · 858 citations
- NL2TL: Transforming Natural Languages to Temporal Logics using Large Language ModelsYongchao Chen, Rujul Gandhi, Yang Zhang, Chuchu FanEMNLP 2023 · 48 citations
- DeepSTL - From English Requirements to Signal Temporal LogicJie He, Ezio Bartocci, Dejan Nickovic, Haris Isakovic et al.ICSE 2022 · 35 citations
- Generating Critical Test Scenarios for Autonomous Driving Systems via Influential Behavior PatternsHaoxiang Tian, Guoquan Wu, Jiren Yan, Yan Jiang et al.ASE 2022 · 24 citations
Related papers
- Bridging Natural Language and Formal Specification-Automated Translation of Software Requirements to LTL via Hierarchical Semantics Decomposition Using LLMsZhi Ma, Cheng Wen, Zhexin Su, Xiao Liang et al.ASE 2025 · 3 citations
- ADARULE: LLM-Driven Natural Language to LTL Conversion via Pattern-Adaptive Rule InductionJiayi Hu, Jingling Sun, Chong Wang, Yihao Huang et al.ICSE 2026
- Modeling Like Peeling an Onion: Layerwise Analysis-Driven Automatic Behavioral Model GenerationYike Huang, Ming Hu, Xiaohong Chen, Zhi Jin et al.ICSE 2026
- VERIFY: A Novel Multi-Domain Dataset Grounding LTL in Contextual Natural Language via Provable Intermediate LogicPaapa Quansah, Pablo Rivas, Ernest BonnahICLR 2026
- Validating Formal Specifications with LLM-Generated Test CasesAlcino Cunha, Nuno MacedoFM 2026 · 1 citation
