Lune

ISSTA2026Top-tier venue

On the Role of Large Language Models in Robustness-Guided Requirement Falsification

Ali Kaya, Ivan Porres

2026Year

Abstract

Robustness-guided falsification supports the validation of cyber-physical systems by searching for system traces that violate formal temporal requirements. Existing falsification methods typically treat candidate generation as a numerical search over the input space, guided by quantitative robustness feedback. This paper studies a different mechanism: a guarded advisory proposer that uses a large language model (LLM) to generate candidate inputs. The LLM-based proposer suggests candidate inputs using the following information: the requirement to validate, the interface of the system under test (SUT), and the history of prior candidate evaluations. Admission gates ensure that only valid candidate inputs are considered for execution. We evaluate this approach on 19 falsification requirements from the ARCH-COMP 2024 CPS benchmark and on five synthetic problems. The results show that LLM-based proposing is sensitive to model choice and context representation. Among the tested variants, GPT-P, which uses a GPT model with the plain context representation, is the most reliable representative variant. On ARCH-COMP, GPT-P reaches perfect falsification on 16 of 19 requirements, obtains at least one falsification on all 19 requirements, and, on 8 requirements, uses the fewest executions among the reported tools that also achieve perfect falsification. However, its performance is heterogeneous. We explain this behavior through three classes of falsification problems: informative context cues, ambiguous context cues with a strong robustness gradient, and ambiguous context cues with a weak robustness gradient. Synthetic benchmark problems reproduce these classes under controlled mechanisms. Overall, GPT-P performs best on the informative-context-cue class, where the requirement and SUT interface expose useful candidate hypotheses that the LLM can exploit before extensive robustness-feedback search is needed. To our knowledge, existing robustness-guided falsification methods do not exploit such semantic context cues when proposing candidate inputs.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get cf472753-d369-4932-abfc-72f18cdc6dca

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines