Understanding the Effectiveness of Coverage Criteria for Large Language Models: A Special Angle from Jailbreak Attacks
Shide Zhou, Tianlin Li, Kailong Wang, Yihao Huang, Ling Shi, Yang Liu, Haoyu Wang
Abstract
Large language models (LLMs) have revolutionized artificial intelligence, but their increasing deployment across critical domains has raised concerns about their abnormal behaviors when faced with malicious attacks. Such vulnerability alerts the widespread inadequacy of pre-release testing. In this paper, we conduct a comprehensive empirical study to evaluate the effectiveness of traditional coverage criteria in identifying such inadequacies, exemplified by the significant security concern of jailbreak attacks. Our study begins with a clustering analysis of the hidden states of LLMs, revealing that the embedded characteristics effectively distinguish between different query types. We then systematically evaluate the performance of these criteria across three key dimensions: criterion level, layer level, and token level. Our research uncovers significant differences in neuron coverage when LLMs process normal versus jailbreak queries, aligning with our clustering experiments. Leveraging these findings, we propose three practical applications of coverage criteria in the context of LLM security testing. Specifically, we develop a realtime jailbreak detection mechanism that achieves high accuracy (93.61 % on average) in classifying queries as normal or jailbreak. Furthermore, we explore the use of coverage levels to prioritize test cases, improving testing efficiency by focusing on high-risk interactions and removing redundant tests. Lastly, we introduce a coverage-guided approach for generating jailbreak attack examples, enabling systematic refinement of prompts to uncover vulnerabilities. This study improves our understanding of LLM security testing, enhances their safety, and provides a foundation for developing more robust AI applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba46a3ad-dd42-46ad-ad75-2b1d45169021Cited by top-tier papers3
- Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM InputsJuyeon Yoon, Somin Kim, Robert Feldt, Shin YooFSE 2026 · 2 citations
- Efficient Input-Level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation VariationShengfang Zhai, Jiajun Li, Yue Liu, Huanran Chen et al.ICCV 2025 · 2 citations
- Validating LLM-Generated SQL Queries through Metamorphic PromptingLi Lin, Qinglin Zhu, Jintai Hong, Chong Wang et al.FSE 2026
Builds on14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Fuzz testing based data augmentation to improve robustness of deep neural networksXiang Gao, Ripon K. Saha, Mukul R. Prasad, Abhik RoychoudhuryICSE 2020 · 116 citations
- Perception-Guided Jailbreak Against Text-to-Image ModelsYihao Huang, Le Liang, Tianlin Li, Xiaojun Jia et al.AAAI 2025 · 34 citations
- ReCode: Robustness Evaluation of Code Generation ModelsShiqi Wang, Zheng Li, Haifeng Qian, Chenghao Yang et al.ACL 2023 · 32 citations
Related papers
- From "Sure" to "Sorry": Detecting Jailbreak in Large Vision Language Model via JailNeuronsYuyou Gan, Qingming Li, Junhao Li, Zhi Chen et al.ICLR 2026
- Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language ModelsLang Gao, Jiahui Geng, Xiangliang Zhang, Preslav Nakov et al.ACL 2025
- From LLMs to MLLMs: Exploring the Landscape of Multimodal JailbreakingSiyuan Wang, Zhuohan Long, Zhihao Fan, Zhongyu WeiEMNLP 2024 · 5 citations
- Smoke and Mirrors: Jailbreaking LLM-based Code Generation via Implicit Malicious PromptsSheng Ouyang, Yihao Qin, Bo Lin, Liqian Chen et al.ICSE 2026
- Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language ModelsZhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron et al.USENIX Security 2024 · 103 citations
