Exploiting Task-Level Vulnerabilities: An Automatic Jailbreak Attack and Defense Benchmarking for LLMs
Lan Zhang, Xinben Gao, Liuyi Yao, Jinke Song, Yaliang Li
摘要
Recent advancements in large language models (LLMs) have notably improved their proficiency in executing complex tasks. However, these advancements are accompanied by an increased risk of generating toxic content as well as leaking private information. "Jailbreak" is an emerging trend to amplify this vulnerability by carefully modifying prompts such as "DAN" to circumvent the LLMs' defense. Notwithstanding, existing jailbreaks typically focus on specific prompts or tokens, rendering them susceptible to countermeasures such as realignments. In contrast to these prompt-level or tokenlevel jailbreaks, we present a novel task-level jailbreak based on "knowledge decomposition" , which does not rely on any specific prompts or tokens. Our attack demonstrates significantly enhanced resistance against realignments compared to previous jailbreak techniques. Furthermore, our attack not only achieves about 10% higher success rates than SOTA attacks but also generates responses that are richer in detail and information. This is attributed to aggregation of responses from multiple well-designed queries rather than relying on only a singular query as in previous attacks, thus signifying an elevated risk of threat. On the other hand, "knowledge decomposition" provide us a method to generate plenty tasks with varying risk levels, thereby establishing a novel benchmark to assess the defensive effectiveness of LLMs. Warning: this paper contains content that can be offensive in nature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
相关 Paper
- TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language ModelsZhi Xu, Jiaqi Li, Xiaotong Zhang, Hong Yu 等ICLR 2026 · 被引用 2 次
- Multi-Turn Jailbreaking Large Language Models via Attention ShiftingXiaohu Du, Fan Mo, Ming Wen, Tu Gu 等AAAI 2025 · 被引用 26 次
- Stand on The Shoulders of Giants: Building JailExpert from Previous Attack ExperienceXi Wang, Songlei Jian, Shasha Li, Xiaopeng Li 等EMNLP 2025 · 被引用 1 次
- Towards Understanding Jailbreak Attacks in LLMs: A Representation Space AnalysisYuping Lin, Pengfei He, Han Xu, Yue Xing 等EMNLP 2024 · 被引用 6 次
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task ConcurrencyYukun Jiang, Mingjie Li, Michael Backes, Yang ZhangNeurIPS 2025 · 被引用 17 次
