D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting
Huanli Gong, Zhipeng Wei, Yu Fu, Haz Shahgir, Ananya Gupta, Yue Dong, N. Benjamin Erichson
Abstract
Multi-turn jailbreak attacks pose a growing threat to large language model (LLM) safety because they exploit feedback from auxiliary judge models to iteratively refine prompts toward harmful goals. Existing defenses largely detect or block unsafe content at individual turns or at the final response, leaving the judge-driven refinement loop intact and allowing attackers to extract informative feedback from intermediate interactions. We introduce D-Judge, a semantics-preserving output rewriting defense that intervenes directly in this loop by rewriting the victim LLM’s responses before they are evaluated by the attacker’s judge. By misaligning the judge’s feedback signal without changing the meaning of the original response, D-Judge derails the attacker’s prompt-refinement process, causing subsequent queries to be optimized against a distorted signal of attack progress. To improve D-Judge’s ability to produce such rewrites, we construct a dataset of semantically equivalent response pairs that induce different judge-assigned harmfulness scores, and use it for supervised fine-tuning followed by direct preference optimization. Experiments on HarmBench show that D-Judge reduces the success rate of state-of-the-art multi-turn jailbreaks while preserving performance on benign benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
- Improving Alignment and Robustness with Circuit BreakersAndy Zou, Long Phan, Justin Wang, Derek Duenas et al.NeurIPS 2024 · 362 citations
Related papers
- MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language ModelsSiyu Yan, Long Zeng, Xuecheng Wu, Chengcheng Han et al.EMNLP 2025 · 1 citation
- MetaDefense: Defending Fine-tuning based Jailbreak Attack Before and During GenerationWeisen Jiang, Sinno Jialin PanNeurIPS 2025 · 10 citations
- MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror CraftingRui Pu, Chaozhuo Li, Rui Ha, Litian Zhang et al.AAAI 2026
- Multi-Turn Jailbreaking Large Language Models via Attention ShiftingXiaohu Du, Fan Mo, Ming Wen, Tu Gu et al.AAAI 2025 · 26 citations
- TROJail: Trajectory-Level Optimization for Multi-Turn Large Language Model Jailbreaks with Process RewardsXiqiao Xiong, Ouxiang Li, Zhuo Liu, Moxin Li et al.ACL 2026 · 7 citations
