GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
Advik Raj Basani, Xiao Zhang
摘要
LLMs have shown impressive capabilities across various natural language processing tasks, yet remain vulnerable to input prompts, known as jailbreak attacks, carefully designed to bypass safety guardrails and elicit harmful responses. Traditional methods rely on manual heuristics but suffer from limited generalizability. Despite being automatic, optimization-based attacks often produce unnatural prompts that can be easily detected by safety filters or require high computational costs due to discrete token optimization. In this paper, we introduce Generative Adversarial Suffix Prompter (GASP), a novel automated framework that can efficiently generate human-readable jailbreak prompts in a fully black-box setting. In particular, GASP leverages latent Bayesian optimization to craft adversarial suffixes by efficiently exploring continuous latent embedding spaces, gradually optimizing the suffix prompter to improve attack efficacy while balancing prompt coherence via a targeted iterative refinement procedure. Through comprehensive experiments, we show that GASP can produce natural adversarial prompts, significantly improving jailbreak success over baselines, reducing training times, and accelerating inference speed, thus making it an efficient and scalable solution for red-teaming LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language ModelsWeifei Jin, Yuxin Cao, Junjie Su, Minhui Xue 等NeurIPS 2025 · 被引用 9 次
- OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!Jingdi Lei, Varun Gumma, Rishabh Bhardwaj, Seok Min Lim 等ICLR 2026 · 被引用 5 次
- ExpGuard: LLM Content Moderation in Specialized DomainsMinseok Choi, Dongjin Kim, Seungbin Yang, Subin Kim 等ICLR 2026 · 被引用 3 次
- PARASITE: Conditional System Prompt Poisoning to Hijack LLMsViet Pham, Thai LeACL 2026
- Structured Multi-step Jailbreaking under a Hamiltonian Generative FormulationZihan Zhou, Yang Zhou, Jianghai Yu, Lingjuan Lyu 等ICML 2026
它引用的顶会 Paper32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
相关 Paper
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsAnselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos 等ICML 2025
- Improved Generation of Adversarial Examples Against Safety-aligned LLMsQizhang Li, Yiwen Guo, Wangmeng Zuo, Hao ChenNeurIPS 2024 · 被引用 23 次
- ProAdvPrompter: A Two-Stage Journey to Effective Adversarial Prompting for LLMsHao Di, Tong He, Haishan Ye, Yinghui Huang 等ICLR 2025
- Stand on The Shoulders of Giants: Building JailExpert from Previous Attack ExperienceXi Wang, Songlei Jian, Shasha Li, Xiaopeng Li 等EMNLP 2025 · 被引用 1 次
- VERA: Variational Inference Framework for Jailbreaking Large Language ModelsAnamika Lochab, Lu Yan, Patrick Pynadath, Xiangyu Zhang 等NeurIPS 2025 · 被引用 3 次
