Adversarial Reasoning at Jailbreaking Time
Mahdi Sabbaghi, Paul Kassianik, George J. Pappas, Amin Karbasi, Hamed Hassani
摘要
As large language models (LLMs) are becoming more capable and widespread, the study of their failure cases is becoming increasingly important. Recent advances in standardizing, measuring, and scaling test-time compute suggest new methodologies for optimizing models to achieve high performance on hard tasks. In this paper, we apply these advances to the task of "model jailbreaking": eliciting harmful responses from aligned LLMs. We develop an adversarial reasoning approach to automatic jailbreaking via test-time computation that achieves SOTA attack success rates (ASR) against many aligned LLMs, even the ones that aim to trade inference-time compute for adversarial robustness. Our approach introduces a new paradigm in understanding LLM vulnerabilities, laying the foundation for the development of more robust and trustworthy AI systems. Code is available at https://github.com/Helloworld10011/Adversarial-Reasoning .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Parallel-R1: Towards Parallel Thinking via Reinforcement LearningTong Zheng, Hongming Zhang, Wenhao Yu, Xiaoyang Wang 等ICLR 2026 · 被引用 53 次
- Antidistillation SamplingYash Savani, Asher Trockman, Zhili Feng, Yixuan Even Xu 等NeurIPS 2025 · 被引用 24 次
- ARMOR: Aligning Secure and Safe Large Language Models via Meticulous ReasoningZhengyue Zhao, Yingzi Ma, Somesh Jha, Marco Pavone 等ICLR 2026 · 被引用 8 次
- One Token Embedding Is Enough to Deadlock Your Large Reasoning ModelMohan Zhang, Yihua Zhang, Jinghan Jia, Zhangyang (Atlas) Wang 等NeurIPS 2025 · 被引用 7 次
- EcoAlign: An Economically Rational Framework for Efficient LVLM AlignmentRuoxi Cheng, Haoxuan Ma, Teng Ma, Hongyi ZhangCVPR 2026 · 被引用 6 次
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
相关 Paper
- Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak AttacksYue Zhou, Henry Peng Zou, Barbara Di Eugenio, Yang ZhangEMNLP 2024 · 被引用 3 次
- Weak-to-Strong Jailbreaking on Large Language ModelsXuandong Zhao, Xianjun Yang, Tianyu Pang, Chao Du 等ICML 2025
- Defending Against Alignment-Breaking Attacks via Robustly Aligned LLMBochuan Cao, Yuanpu Cao, Lu Lin, Jinghui ChenACL 2024 · 被引用 34 次
- from Benign import Toxic: Jailbreaking the Language Model via Adversarial MetaphorsYu Yan, Sheng Sun, Zenghao Duan, Teli Liu 等ACL 2025 · 被引用 14 次
- Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical EvidenceShaopeng Fu, Liang Ding, Jingfeng Zhang, Di WangNeurIPS 2025 · 被引用 15 次
