STAR: Strategy-driven Automatic Jailbreak Red-teaming For Large Language Model
Jianing Liu, Qingming Li, Jiahao Chen, Rui Zeng, Binbin Zhao, Shouling Ji
Abstract
Jailbreaking refers to techniques that bypass the safety alignment of large language models (LLMs) to elicit harmful outputs, and automated red-teaming has become a key approach for detecting such vulnerabilities before deployment. However, most existing red-teaming methods operate directly in text space, where they tend to generate semantically similar prompts and thus fail to probe the broader spectrum of latent vulnerabilities within a model. To address this limitation, we shift the exploration of jailbreaking strategies from conventional text space to the model's latent activation space and propose STAR (STrategy-driven Automatic Jailbreak Red-teaming), a black-box framework for systematically generating jailbreak prompts. STAR is composed of two modules: (i) strategy generation module, which extracts the principal components of existing strategies and recombines them to generate novel ones; and (ii) prompt generation module, which translates abstract strategies into concrete jailbreak prompts with high success rates. Experimental results show that STAR substantially outperforms state-of-the-art baselines in terms of both attack success rate and strategy diversity. These findings highlight critical vulnerabilities in current alignment techniques and establish STAR as a more powerful paradigm for comprehensive LLM security evaluation.
Warning: This paper contains unfiltered and potentially harmful text.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Lifelong Safety Alignment for Language ModelsHaoyu Wang, Yifei Zhao, Zeyu Qin, Chao Du et al.NeurIPS 2025 · 18 citations
- TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy ExplorationChunxiao Li, Lijun Li, Jing ShaoCVPR 2026 · 4 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- CoP: Agentic Red-teaming for Large Language Models using Composition of PrinciplesChen Xiong, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2025 · 13 citations
- Distract Large Language Models for Automatic Jailbreak AttackZeguan Xiao, Yan Yang, Guanhua Chen, Yun ChenEMNLP 2024 · 8 citations
- Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language ModelsYanjiang Liu, Shuheng Zhou, Yaojie Lu, Huijia Zhu et al.ICLR 2026 · 10 citations
- Towards Understanding Jailbreak Attacks in LLMs: A Representation Space AnalysisYuping Lin, Pengfei He, Han Xu, Yue Xing et al.EMNLP 2024 · 6 citations
- Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language ModelsKai Hu, Abhinav Aggarwal, Mehran Khodabandeh, David Zhang et al.ACL 2026
