SHARP: Self-adaptive Harmful Category-aware Prompt Generation for Black-box Jailbreaking
Yingjie Xue, Xingyou Xia, Jun Zhang, Yunbo Cao, Dengpan Ye, Guotong Geng, Fei Li
Abstract
Large Language Models (LLMs) have been widely applied in various domains such as education and healthcare, making safety assurance crucial. Jailbreak attacks, a method used in red-teaming, can help evaluate and improve the defensive strategies of LLMs. However, existing jailbreak methods often overlook the semantic differences across categories of harmful questions, leading to inconsistent success rates and reduced overall attack effectiveness. We propose the first category-aware jailbreak framework, SHARP, which incorporates the semantic category of harmful questions into prompt generation. Trained on a verified jailbreak dataset, SHARP enables the model to learn category-specific semantic features and adaptively generate prompts that bypass safety mechanisms. The method combines two-stage LoRA fine-tuning, and DPO-based reinforcement learning to optimize both attack success and category alignment. Experiments show that SHARP significantly improves attack success rates and achieves better cross-category robustness compared to the state-of-the-art (SOTA) baselines, providing an efficient and scalable tool for evaluating LLM safety.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca7a5efc-ba79-4afd-816f-9f412da352f4Builds on11
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji et al.ICLR 2024 · 656 citations
- Fine-Grained Human Feedback Gives Better Rewards for Language Model TrainingZeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri et al.NeurIPS 2023 · 516 citations
- Hard Prompts Made Easy: Gradient-Based Discrete Optimization for Prompt Tuning and DiscoveryYuxin Wen, Neel Jain, John Kirchenbauer, Micah Goldblum et al.NeurIPS 2023 · 454 citations
Related papers
- Alignment-Weighted DPO: A principled reasoning approach to improve safety alignmentMengxuan Hu, Vivek V. Datla, Anoop Kumar, Zihan Guan et al.ICLR 2026 · 3 citations
- TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language ModelsZhi Xu, Jiaqi Li, Xiaotong Zhang, Hong Yu et al.ICLR 2026 · 2 citations
- Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?Sravanti Addepalli, Yerram Varun, Arun Suggala, Karthikeyan Shanmugam et al.ICLR 2025
- Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language ModelsKai Hu, Abhinav Aggarwal, Mehran Khodabandeh, David Zhang et al.ACL 2026
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsAnselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos et al.ICML 2025
