Lune

ISSTA2026顶会

On the Role of Large Language Models in Robustness-Guided Requirement Falsification

Ali Kaya, Ivan Porres

2026年份

摘要

Robustness-guided falsification supports the validation of cyber-physical systems by searching for system traces that violate formal temporal requirements. Existing falsification methods typically treat candidate generation as a numerical search over the input space, guided by quantitative robustness feedback. This paper studies a different mechanism: a guarded advisory proposer that uses a large language model (LLM) to generate candidate inputs. The LLM-based proposer suggests candidate inputs using the following information: the requirement to validate, the interface of the system under test (SUT), and the history of prior candidate evaluations. Admission gates ensure that only valid candidate inputs are considered for execution. We evaluate this approach on 19 falsification requirements from the ARCH-COMP 2024 CPS benchmark and on five synthetic problems. The results show that LLM-based proposing is sensitive to model choice and context representation. Among the tested variants, GPT-P, which uses a GPT model with the plain context representation, is the most reliable representative variant. On ARCH-COMP, GPT-P reaches perfect falsification on 16 of 19 requirements, obtains at least one falsification on all 19 requirements, and, on 8 requirements, uses the fewest executions among the reported tools that also achieve perfect falsification. However, its performance is heterogeneous. We explain this behavior through three classes of falsification problems: informative context cues, ambiguous context cues with a strong robustness gradient, and ambiguous context cues with a weak robustness gradient. Synthetic benchmark problems reproduce these classes under controlled mechanisms. Overall, GPT-P performs best on the informative-context-cue class, where the requirement and SUT interface expose useful candidate hypotheses that the LLM can exploit before extensive robustness-feedback search is needed. To our knowledge, existing robustness-guided falsification methods do not exploit such semantic context cues when proposing candidate inputs.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖