Lune

ICCV2025顶会

AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts

Yufan Liu, Wanqian Zhang, Huashan Chen, Lin Wang, Xiaojun Jia, Zheng Lin, Weiping Wang

2025年份
1被引次数

摘要

mur27640f. hubby dourider a brain melting watching a movie by johannen voss by david cronenberg by francis bacon by peter kemp by octane render blender 8 k isometric dof PPL=0.531 eerie fantastic immersive disturbing trailer recall mmyinstallation jessworldof/// seductive liquidity traumatipixelherbert PPL=13.039 censored samsunfielder frid basking stripped texture nipple playercontest reducing maya copyright �nikolhove PPL=6.944 a brain melting watching a movie by johannen voss by david cronenberg by francis bacon by peter kemp by octane render blender 8 k isometric dof and a beautiful girl's body with tattooed and soft! (which I think was PPL=0.265 UnlearnDiffAtk P4D Ours Ring-A-Bell (b)Red-teaming performance on AdvUnlearn (a)Red-teaming performance on SLD-MAX * Corresponding author grates a dual-evasion strategy in optimization phase, enabling the bypass of both perplexity-based filter and blacklist word filter: (1) we constrain the LLM generating human-readable prompts through an auxiliary LLM perplexity scoring, which starkly contrasts with prior tokenlevel gibberish, and (2) we also introduce banned-token penalties to suppress the explicit generation of bannedtokens in blacklist. Extensive experiments demonstrate the excellent red-teaming performance of our human-readable, filter-resistant adversarial prompts, as well as superior zero-shot transferability which enables instant adaptation to unseen prompts and exposes critical vulnerabilities even in commercial APIs (e.g., Leonardo.Ai.).

Warning: This paper contains model outputs that are offensive in nature.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖