Adversarial Reasoning at Jailbreaking Time
Mahdi Sabbaghi, Paul Kassianik, George J. Pappas, Amin Karbasi, Hamed Hassani
Abstract
As large language models (LLMs) are becoming more capable and widespread, the study of their failure cases is becoming increasingly important. Recent advances in standardizing, measuring, and scaling test-time compute suggest new methodologies for optimizing models to achieve high performance on hard tasks. In this paper, we apply these advances to the task of "model jailbreaking": eliciting harmful responses from aligned LLMs. We develop an adversarial reasoning approach to automatic jailbreaking via test-time computation that achieves SOTA attack success rates (ASR) against many aligned LLMs, even the ones that aim to trade inference-time compute for adversarial robustness. Our approach introduces a new paradigm in understanding LLM vulnerabilities, laying the foundation for the development of more robust and trustworthy AI systems. Code is available at https://github.com/Helloworld10011/Adversarial-Reasoning .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Parallel-R1: Towards Parallel Thinking via Reinforcement LearningTong Zheng, Hongming Zhang, Wenhao Yu, Xiaoyang Wang et al.ICLR 2026 · 53 citations
- Antidistillation SamplingYash Savani, Asher Trockman, Zhili Feng, Yixuan Even Xu et al.NeurIPS 2025 · 24 citations
- ARMOR: Aligning Secure and Safe Large Language Models via Meticulous ReasoningZhengyue Zhao, Yingzi Ma, Somesh Jha, Marco Pavone et al.ICLR 2026 · 8 citations
- One Token Embedding Is Enough to Deadlock Your Large Reasoning ModelMohan Zhang, Yihua Zhang, Jinghan Jia, Zhangyang (Atlas) Wang et al.NeurIPS 2025 · 7 citations
- EcoAlign: An Economically Rational Framework for Efficient LVLM AlignmentRuoxi Cheng, Haoxuan Ma, Teng Ma, Hongyi ZhangCVPR 2026 · 6 citations
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson et al.NeurIPS 2024 · 835 citations
Related papers
- Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak AttacksYue Zhou, Henry Peng Zou, Barbara Di Eugenio, Yang ZhangEMNLP 2024 · 3 citations
- Weak-to-Strong Jailbreaking on Large Language ModelsXuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du et al.ICML 2025
- Defending Against Alignment-Breaking Attacks via Robustly Aligned LLMBochuan Cao, Yuanpu Cao, Lu Lin, Jinghui ChenACL 2024 · 34 citations
- from Benign import Toxic: Jailbreaking the Language Model via Adversarial MetaphorsYu Yan, Sheng Sun, Zenghao Duan, Teli Liu et al.ACL 2025 · 14 citations
- Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical EvidenceShaopeng Fu, Liang Ding, Jingfeng Zhang, Di WangNeurIPS 2025 · 15 citations
