ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
Hao Wang, Hao Li, Minlie Huang, Lei Sha
Abstract
The safety defense methods of Large language models (LLMs) stays limited because the dangerous prompts are manually curated to just few known attack types, which fails to keep pace with emerging varieties. Recent studies found that attaching suffixes to harmful instructions can hack the defense of LLMs and lead to dangerous outputs. However, similar to traditional text adversarial attacks, this approach, while effective, is limited by the challenge of the discrete tokens. This gradient based discrete optimization attack requires over 100,000 LLM calls, and due to the unreadable of adversarial suffixes, it can be relatively easily penetrated by common defense methods such as perplexity filters. To cope with this challenge, in this paper, we propose an Adversarial Suffix Embedding Translation Framework (ASETF), aimed at transforming continuous adversarial suffix embeddings into coherent and understandable text. This method greatly reduces the computational overhead during the attack process and helps to automatically generate multiple adversarial samples, which can be used as data to strengthen LLM's security defense. Experimental evaluations were conducted on Llama2, Vicuna, and other prominent LLMs, employing harmful directives sourced from the Advbench dataset. The results indicate that our method significantly reduces the computation time of adversarial suffixes and achieves a much better attack success rate than existing techniques, while significantly enhancing the textual fluency of the prompts. In addition, our approach can be generalized into a broader method for generating transferable adversarial suffixes that can successfully attack multiple LLMs, even black-box LLMs, such as ChatGPT and Gemini.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3fa3ac8-e964-4584-b790-6bfb1d5fd757Cited by top-tier papers9
- LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution ShiftsQibing Ren, Hao Li, Dongrui Liu, Zhanxu Xie et al.ACL 2025 · 53 citations
- GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak MethodsRuixuan Huang, Xunguang Wang, Zongjie Li, Daoyuan Wu et al.ICLR 2026 · 11 citations
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM HallucinationsBuyun Liang, Liangzu Peng, Jinqi Luo, Darshan Thaker et al.NeurIPS 2025 · 11 citations
- Response Attack: Exploiting Contextual Priming to Jailbreak Large Language ModelsZiqi Miao, Lijun Li, Yuan Xiong, Zhenhua Liu et al.AAAI 2026 · 8 citations
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 6 citations
Builds on6
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- Gradient-guided Unsupervised Lexically Constrained Text GenerationLei ShaEMNLP 2020 · 35 citations
Related papers
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsAnselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos et al.ICML 2025
- Improved Generation of Adversarial Examples Against Safety-aligned LLMsQizhang Li, Yiwen Guo, Wangmeng Zuo, Hao ChenNeurIPS 2024 · 23 citations
- GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMsAdvik Raj Basani, Xiao ZhangNeurIPS 2025 · 16 citations
- ProAdvPrompter: A Two-Stage Journey to Effective Adversarial Prompting for LLMsHao Di, Tong He, Haishan Ye, Yinghui Huang et al.ICLR 2025
- Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical EvidenceShaopeng Fu, Liang Ding, Jingfeng Zhang, Di WangNeurIPS 2025 · 15 citations
