Automating Requirements Formalization: Using LLMs and Low-Complexity Distinguishing Traces for Semantic Validation
Daniel Mendoza, Anastasia Mavridou, Andreas Katis, Caroline Trippel
摘要
Translating natural language (NL) requirements into formal specifications is critical for verifying safety-critical systems, but it is error-prone and time-consuming when done manually. While Large Language Models (LLMs) can automate this translation, they often produce incorrect outputs that require extensive validation. In this paper, we propose ARTEMIS, an LLM-based framework that translates unstructured NL requirements into formal temporal logic (TL) specifications. Our framework reduces validation effort through three synergistic, automated techniques: (i) LLM Translation to Structured NL: We use LLMs to translate unstructured NL requirements into structured NL, which has an unambiguous mapping to TL. This intermediate representation reduces translation errors. (ii) Sub-Specification Generation: We generate low-complexity execution traces (i.e., system behaviors) that correspond to candidate specification fragments from the LLM translations. Users inspect these and accept or reject them. (iii) Balanced Distinguishing Trace Generation: We minimize the number of traces users need to inspect by pruning the candidate specification space. Each accepted or rejected trace eliminates candidates logarithmically. We evaluate ARTEMIS on five real-world safety-critical requirements datasets. The results show that it achieves 1.57X higher translation accuracy while reducing manual validation effort by up to 10.83X compared to state-of-the-art baselines.
• Software and its engineering → Formal methods; Requirements analysis; • Computing methodologies → Natural language processing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models Cannot Self-Correct Reasoning YetJie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng 等ICLR 2024 · 被引用 858 次
- NL2TL: Transforming Natural Languages to Temporal Logics using Large Language ModelsYongchao Chen, Rujul Gandhi, Yang Zhang, Chuchu FanEMNLP 2023 · 被引用 48 次
- DeepSTL - From English Requirements to Signal Temporal LogicJie He, Ezio Bartocci, Dejan Nickovic, Haris Isakovic 等ICSE 2022 · 被引用 35 次
- Generating Critical Test Scenarios for Autonomous Driving Systems via Influential Behavior PatternsHaoxiang Tian, Guoquan Wu, Jiren Yan, Yan Jiang 等ASE 2022 · 被引用 24 次
相关 Paper
- Bridging Natural Language and Formal Specification-Automated Translation of Software Requirements to LTL via Hierarchical Semantics Decomposition Using LLMsZhi Ma, Cheng Wen, Zhexin Su, Xiao Liang 等ASE 2025 · 被引用 3 次
- ADARULE: LLM-Driven Natural Language to LTL Conversion via Pattern-Adaptive Rule InductionJiayi Hu, Jingling Sun, Chong Wang, Yihao Huang 等ICSE 2026
- Modeling Like Peeling an Onion: Layerwise Analysis-Driven Automatic Behavioral Model GenerationYike Huang, Ming Hu, Xiaohong Chen, Zhi Jin 等ICSE 2026
- VERIFY: A Novel Multi-Domain Dataset Grounding LTL in Contextual Natural Language via Provable Intermediate LogicPaapa Quansah, Pablo Rivas, Ernest BonnahICLR 2026
- Validating Formal Specifications with LLM-Generated Test CasesAlcino Cunha, Nuno MacedoFM 2026 · 被引用 1 次
