On the Role of Large Language Models in Robustness-Guided Requirement Falsification
Ali Kaya, Ivan Porres
摘要
Robustness-guided falsification supports the validation of cyber-physical systems by searching for system traces that violate formal temporal requirements. Existing falsification methods typically treat candidate generation as a numerical search over the input space, guided by quantitative robustness feedback. This paper studies a different mechanism: a guarded advisory proposer that uses a large language model (LLM) to generate candidate inputs. The LLM-based proposer suggests candidate inputs using the following information: the requirement to validate, the interface of the system under test (SUT), and the history of prior candidate evaluations. Admission gates ensure that only valid candidate inputs are considered for execution. We evaluate this approach on 19 falsification requirements from the ARCH-COMP 2024 CPS benchmark and on five synthetic problems. The results show that LLM-based proposing is sensitive to model choice and context representation. Among the tested variants, GPT-P, which uses a GPT model with the plain context representation, is the most reliable representative variant. On ARCH-COMP, GPT-P reaches perfect falsification on 16 of 19 requirements, obtains at least one falsification on all 19 requirements, and, on 8 requirements, uses the fewest executions among the reported tools that also achieve perfect falsification. However, its performance is heterogeneous. We explain this behavior through three classes of falsification problems: informative context cues, ambiguous context cues with a strong robustness gradient, and ambiguous context cues with a weak robustness gradient. Synthetic benchmark problems reproduce these classes under controlled mechanisms. Overall, GPT-P performs best on the informative-context-cue class, where the requirement and SUT interface expose useful candidate hypotheses that the LLM can exploit before extensive robustness-feedback search is needed. To our knowledge, existing robustness-guided falsification methods do not exploit such semantic context cues when proposing candidate inputs.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Scenario-Based Flexible Modeling and Scalable Falsification for Reconfigurable CPSsJiawan Wang, Wenxia Liu, Muzimiao Zhang, Jiaqi Wei 等CAV 2024 · 被引用 3 次
- Effective Hybrid System Falsification Using Monte Carlo Tree Search Guided by QB-RobustnessZhenya Zhang, Deyun Lyu, Paolo Arcaini, Lei Ma 等CAV 2021 · 被引用 39 次
- Parametric Falsification of Many Probabilistic Requirements Under FlakinessMatteo Camilli, Raffaela MirandolaICSE 2025
- A New Framework for Cybersecurity Refusals in AI AgentsEliot Jones, Matt Fredrikson, Zico KolterICML 2026
- An LLM can Fool Itself: A Prompt-Based Adversarial AttackXilie Xu, Keyi Kong, Ning Liu, Lizhen Cui 等ICLR 2024 · 被引用 146 次
