Lune

USENIX Security2025顶会

Exploiting Task-Level Vulnerabilities: An Automatic Jailbreak Attack and Defense Benchmarking for LLMs

Lan Zhang, Xinben Gao, Liuyi Yao, Jinke Song, Yaliang Li

出版方
2025年份

摘要

Recent advancements in large language models (LLMs) have notably improved their proficiency in executing complex tasks. However, these advancements are accompanied by an increased risk of generating toxic content as well as leaking private information. "Jailbreak" is an emerging trend to amplify this vulnerability by carefully modifying prompts such as "DAN" to circumvent the LLMs' defense. Notwithstanding, existing jailbreaks typically focus on specific prompts or tokens, rendering them susceptible to countermeasures such as realignments. In contrast to these prompt-level or tokenlevel jailbreaks, we present a novel task-level jailbreak based on "knowledge decomposition" , which does not rely on any specific prompts or tokens. Our attack demonstrates significantly enhanced resistance against realignments compared to previous jailbreak techniques. Furthermore, our attack not only achieves about 10% higher success rates than SOTA attacks but also generates responses that are richer in detail and information. This is attributed to aggregation of responses from multiple well-designed queries rather than relying on only a singular query as in previous attacks, thus signifying an elevated risk of threat. On the other hand, "knowledge decomposition" provide us a method to generate plenty tasks with varying risk levels, thereby establishing a novel benchmark to assess the defensive effectiveness of LLMs. Warning: this paper contains content that can be offensive in nature.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext a76e17ec-ee74-4464-a4e6-85ddde3bb00a

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖