Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints
Junxiao Yang, Zhexin Zhang, Shiyao Cui, Hongning Wang, Minlie Huang
Abstract
Jailbreaking attacks can effectively induce unsafe behaviors in Large Language Models (LLMs); however, the transferability of these attacks across different models remains limited. This study aims to understand and enhance the transferability of gradient-based jailbreaking methods, which are among the standard approaches for attacking white-box models. Through a detailed analysis of the optimization process, we introduce a novel conceptual framework to elucidate transferability and identify superfluous constraints-specifically, the response pattern constraint and the token tail constraint-as significant barriers to improved transferability. Removing these unnecessary constraints substantially enhances the transferability and controllability of gradient-based attacks. Evaluated on Llama-3-8B-Instruct as the source model, our method increases the overall Transfer Attack Success Rate (T-ASR) across a set of target models with varying safety levels from 18.4% to 50.3%, while also improving the stability and controllability of jailbreak behaviors on both source and target models. Our code is available at https: //github.com/thu-coai/TransferAttack .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67f8a273-e291-4cd7-bc7e-6d805fa71783Cited by top-tier papers6
- AdvPrefix: An Objective for Nuanced LLM JailbreaksSicheng Zhu, Brandon Amos, Yuandong Tian, Chuan Guo et al.NeurIPS 2025 · 25 citations
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEctionRunqi Lin, Alasdair Paren, Suqin Yuan, Muyang Li et al.CVPR 2026 · 13 citations
- When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' ToxicityShiyao Cui, Xijia Feng, Yingkang Wang, Junxiao Yang et al.AAAI 2026
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMsYixin Tan, Yu Zhe, Rui Wen, Jun SakumaCCS 2026
- Enhancing the Transferability of Jailbreak Attacks on Large Language Models via Exploiting Reparameterization InvarianceAo Wang, Xinghao Yang, Yongshun Gong, Wei Liu et al.ACL 2026
Builds on10
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- AdvPrefix: An Objective for Nuanced LLM JailbreaksSicheng Zhu, Brandon Amos, Yuandong Tian, Chuan Guo et al.NeurIPS 2025 · 25 citations
Related papers
- Attention Eclipse: Manipulating Attention to Bypass LLM Safety-AlignmentPedram Zaree, Md Abdullah Al Mamun, Quazi Mishkatul Alam, Yue Dong et al.EMNLP 2025
- Improved Generation of Adversarial Examples Against Safety-aligned LLMsQizhang Li, Yiwen Guo, Wangmeng Zuo, Hao ChenNeurIPS 2024 · 23 citations
- TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language ModelsZhi Xu, Jiaqi Li, Xiaotong Zhang, Hong Yu et al.ICLR 2026 · 2 citations
- Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained OptimizationKai Hu, Weichen Yu, Yining Li, Tianjun Yao et al.NeurIPS 2024 · 31 citations
- Analogy-based Multi-Turn Jailbreak against Large Language ModelsMengjie Wu, Yihao Huang, Zhenjun Lin, Kangjie Chen et al.NeurIPS 2025 · 9 citations
