LLM-Fuzzer: Scaling Assessment of Large Language Model Jailbreaks
Jiahao Yu, Xingwei Lin, Zheng Yu, Xinyu Xing
摘要
Warning: This paper contains unfiltered content generated by LLMs that may be offensive to readers.
The jailbreak threat poses a significant concern for Large Language Models (LLMs), primarily due to their potential to generate content at scale. If not properly controlled, LLMs can be exploited to produce undesirable outcomes, including the dissemination of misinformation, offensive content, and other forms of harmful or unethical behavior. To tackle this pressing issue, researchers and developers often rely on redteam efforts to manually create adversarial inputs and prompts designed to push LLMs into generating harmful, biased, or inappropriate content. However, this approach encounters serious scalability challenges.
To address these scalability issues, we introduce an automated solution for large-scale LLM jailbreak susceptibility assessment called LLM-FUZZER. Inspired by fuzz testing, LLM-FUZZER uses human-crafted jailbreak prompts as starting points. By employing carefully customized seed selection strategies and mutation mechanisms, LLM-FUZZER generates additional jailbreak prompts tailored to specific LLMs. Our experiments show that LLM-FUZZER-generated jailbreak prompts demonstrate significantly increased effectiveness and transferability. This highlights that many opensource and commercial LLMs suffer from severe jailbreak issues, even after safety fine-tuning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Sok: Evaluating Jailbreak Guardrails for Large Language ModelsXunguang Wang, Zhenlan Ji, Wenxuan Wang, Zongjie Li 等S&P 2026 · 被引用 27 次
- NeuroStrike: Neuron-Level Attacks on Aligned LLMsLichao Wu, Sasha Behrouzi, Mohamadreza Rostami, Maximilian Thang 等NDSS 2026 · 被引用 21 次
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller 等NDSS 2026 · 被引用 17 次
- USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language ModelsBaolin Zheng, Guanlin Chen, Qingyang Teng, Hongqiong Zhong 等ACL 2026 · 被引用 10 次
- TAI3: Testing Agent Integrity in Interpreting User IntentShiwei Feng, Xiangzhe Xu, Xuan Chen, Kaiyuan Zhang 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
相关 Paper
- Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-to-Image Generation ModelsYingkai Dong, Xiangtao Meng, Ning Yu, Zheng Li 等S&P 2025
- LLMs Caught in the Crossfire: Malware Requests and Jailbreak ChallengesHaoyang Li, Huan Gao, Zhiyuan Zhao, Zhiyu Lin 等ACL 2025
- Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language ModelsZhiyuan Yu, Xiaogeng Liu, Shunning Liang, Zach Cameron 等USENIX Security 2024 · 被引用 103 次
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsAnselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos 等ICML 2025
- Efficient LLM-Jailbreaking via Multimodal-LLM JailbreakHaoxuan Ji, Zheng Lin, Zhenxing Niu, Xinbo Gao 等AAAI 2026 · 被引用 4 次
