Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
Hongfu Liu, Yuxi Xie, Ye Wang, Michael Shieh
摘要
Language Language Models (LLMs) face safety concerns due to potential misuse by malicious users. Recent red-teaming efforts have identified adversarial suffixes capable of jailbreaking LLMs using the gradient-based search algorithm Greedy Coordinate Gradient (GCG). However, GCG struggles with computational inefficiency, limiting further investigations regarding suffix transferability and scalability across models and data. In this work, we bridge the connection between search efficiency and suffix transferability. We propose a two-stage transfer learning framework, DeGCG, which decouples the search process into behavioragnostic pre-searching and behavior-relevant post-searching. Specifically, we employ direct first target token optimization in pre-searching to facilitate the search process. We apply our approach to cross-model, cross-data, and self-transfer scenarios. Furthermore, we introduce an interleaved variant of our approach, i-DeGCG, which iteratively leverages selftransferability to accelerate the search process. Experiments on HarmBench demonstrate the efficiency of our approach across various models and domains. Notably, our i-DeGCG outperforms the baseline on Llama2-chat-7b with ASRs of 43.9 (+22.2) and 39.0 (+19.5) on valid and test sets, respectively. Further analysis on cross-model transfer indicates the pivotal role of first target token optimization in leveraging suffix transferability for efficient searching 1 . * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- AdvPrefix: An Objective for Nuanced LLM JailbreaksSicheng Zhu, Brandon Amos, Yuandong Tian, Chuan Guo 等NeurIPS 2025 · 被引用 25 次
- Enhancing the Transferability of Jailbreak Attacks on Large Language Models via Exploiting Reparameterization InvarianceAo Wang, Xinghao Yang, Yongshun Gong, Wei Liu 等ACL 2026
- Structured Multi-step Jailbreaking under a Hamiltonian Generative FormulationZihan Zhou, Yang Zhou, Jianghai Yu, Lingjuan Lyu 等ICML 2026
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li 等ICLR 2024 · 被引用 481 次
相关 Paper
- Lookahead-GCG: Improving Universal Multi-Model Optimization-Based Jailbreaking Attacks via Stochastic Nesterov OptimizationRong Feng, Haohan Zhao, Shiqin Tang, Geng Liu 等ICML 2026
- Improved Techniques for Optimization-Based Jailbreaking on Large Language ModelsXiaojun Jia, Tianyu Pang, Chao Du, Yihao Huang 等ICLR 2025
- Improved Generation of Adversarial Examples Against Safety-aligned LLMsQizhang Li, Yiwen Guo, Wangmeng Zuo, Hao ChenNeurIPS 2024 · 被引用 23 次
- Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous ConstraintsJunxiao Yang, Zhexin Zhang, Shiyao Cui, Hongning Wang 等ACL 2025 · 被引用 6 次
- ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix EmbeddingsHao Wang, Hao Li, Minlie Huang, Lei ShaEMNLP 2024 · 被引用 7 次
