Fight Back Against Jailbreaking via Prompt Adversarial Tuning
Yichuan Mo, Yuji Wang, Zeming Wei, Yisen Wang
摘要
While Large Language Models (LLMs) have achieved tremendous success in various applications, they are also susceptible to jailbreaking attacks. Several primary defense strategies have been proposed to protect LLMs from producing harmful information, mostly focusing on model fine-tuning or heuristical defense designs. However, how to achieve intrinsic robustness through prompt optimization remains an open problem. In this paper, motivated by adversarial training paradigms for achieving reliable robustness, we propose an approach named Prompt Adversarial Tuning (PAT) that trains a prompt control attached to the user prompt as a guard prefix. To achieve our defense goal whilst maintaining natural performance, we optimize the control prompt with both adversarial and benign prompts. Comprehensive experiments show that our method is effective against both grey-box and black-box attacks, reducing the success rate of advanced attacks to nearly 0%, while maintaining the model's utility on the benign task and incurring only negligible computational overhead, charting a new perspective for future explorations in LLM security. Our code is available at https://github.com/PKU-ML/PAT.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Robust Prompt Optimization for Defending Language Models Against Jailbreaking AttacksAndy Zhou, Bo Li, Haohan WangNeurIPS 2024 · 被引用 198 次
- A Theoretical Understanding of Self-Correction through In-context AlignmentYifei Wang, Yuyang Wu, Zeming Wei, Stefanie Jegelka 等NeurIPS 2024 · 被引用 69 次
- Jailbreak Large Vision-Language Models Through Multi-Modal LinkageYu Wang, Xiaofei Zhou, Yichen Wang, Geyuan Zhang 等ACL 2025 · 被引用 51 次
- Sok: Evaluating Jailbreak Guardrails for Large Language ModelsXunguang Wang, Zhenlan Ji, Wenxuan Wang, Zongjie Li 等S&P 2026 · 被引用 27 次
- AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language ModelsKai Li, Can Shen, Yile Liu, Jirui Han 等ICLR 2026 · 被引用 17 次
它引用的顶会 Paper6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey 等ICLR 2020 · 被引用 829 次
- When Adversarial Training Meets Vision Transformers: Recipes from Training to ArchitectureYichuan Mo, Dongxian Wu, Yifei Wang, Yiwen Guo 等NeurIPS 2022 · 被引用 109 次
- Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial AttacksXinyu Zhang, Hanbin Hong, Yuan Hong, Peng Huang 等S&P 2024 · 被引用 41 次
相关 Paper
- Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical EvidenceShaopeng Fu, Liang Ding, Jingfeng Zhang, Di WangNeurIPS 2025 · 被引用 15 次
- Defending Against Alignment-Breaking Attacks via Robustly Aligned LLMBochuan Cao, Yuanpu Cao, Lu Lin, Jinghui ChenACL 2024 · 被引用 34 次
- Dynamic Deep Prompt Optimization for Defending Against Jailbreak Attacks on LLMsDoniyorkhon Obidov, Honggang Yu, Xiaolong Guo, Kaichen YangAAAI 2026
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsAnselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos 等ICML 2025
- MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror CraftingRui Pu, Chaozhuo Li, Rui Ha, Litian Zhang 等AAAI 2026
