On the Role of Large Language Models in Robustness-Guided Requirement Falsification
Ali Kaya, Ivan Porres
Abstract
Robustness-guided falsification supports the validation of cyber-physical systems by searching for system traces that violate formal temporal requirements. Existing falsification methods typically treat candidate generation as a numerical search over the input space, guided by quantitative robustness feedback. This paper studies a different mechanism: a guarded advisory proposer that uses a large language model (LLM) to generate candidate inputs. The LLM-based proposer suggests candidate inputs using the following information: the requirement to validate, the interface of the system under test (SUT), and the history of prior candidate evaluations. Admission gates ensure that only valid candidate inputs are considered for execution. We evaluate this approach on 19 falsification requirements from the ARCH-COMP 2024 CPS benchmark and on five synthetic problems. The results show that LLM-based proposing is sensitive to model choice and context representation. Among the tested variants, GPT-P, which uses a GPT model with the plain context representation, is the most reliable representative variant. On ARCH-COMP, GPT-P reaches perfect falsification on 16 of 19 requirements, obtains at least one falsification on all 19 requirements, and, on 8 requirements, uses the fewest executions among the reported tools that also achieve perfect falsification. However, its performance is heterogeneous. We explain this behavior through three classes of falsification problems: informative context cues, ambiguous context cues with a strong robustness gradient, and ambiguous context cues with a weak robustness gradient. Synthetic benchmark problems reproduce these classes under controlled mechanisms. Overall, GPT-P performs best on the informative-context-cue class, where the requirement and SUT interface expose useful candidate hypotheses that the LLM can exploit before extensive robustness-feedback search is needed. To our knowledge, existing robustness-guided falsification methods do not exploit such semantic context cues when proposing candidate inputs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get cf472753-d369-4932-abfc-72f18cdc6dcaRelated papers
- Scenario-Based Flexible Modeling and Scalable Falsification for Reconfigurable CPSsJiawan Wang, Wenxia Liu, Muzimiao Zhang, Jiaqi Wei et al.CAV 2024 · 3 citations
- Effective Hybrid System Falsification Using Monte Carlo Tree Search Guided by QB-RobustnessZhenya Zhang, Deyun Lyu, Paolo Arcaini, Lei Ma et al.CAV 2021 · 39 citations
- Parametric Falsification of Many Probabilistic Requirements Under FlakinessMatteo Camilli, Raffaela MirandolaICSE 2025
- A New Framework for Cybersecurity Refusals in AI AgentsEliot Jones, Matt Fredrikson, Zico KolterICML 2026
- An LLM can Fool Itself: A Prompt-Based Adversarial AttackXilie Xu, Keyi Kong, Ning Liu, Lizhen Cui et al.ICLR 2024 · 146 citations
