One-Shot is Enough: Consolidating Multi-Turn Attacks into Efficient Single-Turn Prompts for LLMs
Junwoo Ha, Hyunjun Kim, Sangyoon Yu, Haon Park, Ashkan Yousefpour, Yuna Park, Suhyun Kim
摘要
We introduce a novel framework for consolidating multi-turn adversarial "jailbreak" prompts into single-turn queries, significantly reducing the manual overhead required for adversarial testing of large language models (LLMs). While multi-turn human jailbreaks have been shown to yield high attack success rates (ASRs), they demand considerable human effort and time. Our proposed Multi-turn-to-Single-turn (M2S) methods-HYPHENIZE, NUMBERIZE, and PYTHONIZEsystematically reformat multi-turn dialogues into structured single-turn prompts. Despite eliminating iterative back-and-forth interactions, these reformatted prompts preserve and often enhance adversarial potency: in extensive evaluations on the Multi-turn Human Jailbreak (MHJ) dataset, M2S methods yield ASRs ranging from 70.6% to 95.9% across various stateof-the-art LLMs. Remarkably, our single-turn prompts outperform the original multi-turn attacks by up to 17.5% in absolute ASR, while reducing token usage by more than half on average. Further analyses reveal that embedding malicious requests in enumerated or codelike structures exploits "contextual blindness," undermining both native guardrails and external input-output safeguards. By consolidating multi-turn conversations into efficient singleturn prompts, our M2S framework provides a powerful tool for large-scale red-teaming and exposes critical vulnerabilities in contemporary LLM defenses. All code, data, and conversion prompts are available for reproducibility and further investigations: https: //github.com/Junuha/M2S_DATA
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- Improving Alignment and Robustness with Circuit BreakersAndy Zou, Long Phan, Justin Wang, Derek Duenas 等NeurIPS 2024 · 被引用 362 次
- Pretraining Language Models with Human PreferencesTomasz Korbak, Kejian Shi, Angelica Chen, Rasika Vinayak Bhalerao 等ICML 2023 · 被引用 287 次
相关 Paper
- AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsAnselm Paulus, Arman Zharmagambetov, Chuan Guo, Brandon Amos 等ICML 2025
- MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM SafetyJialin Song, Xiaodong Liu, Weiwei Yang, Wuyang Chen 等ICML 2026 · 被引用 5 次
- DAMON: A Dialogue-Aware MCTS Framework for Jailbreaking Large Language ModelsXu Zhang, Xunjian Yin, Dinghao Jing, Huixuan Zhang 等EMNLP 2025 · 被引用 2 次
- Multi-Turn Jailbreaking Large Language Models via Attention ShiftingXiaohu Du, Fan Mo, Ming Wen, Tu Gu 等AAAI 2025 · 被引用 26 次
- Analogy-based Multi-Turn Jailbreak against Large Language ModelsMengjie Wu, Yihao Huang, Zhenjun Lin, Kangjie Chen 等NeurIPS 2025 · 被引用 9 次
