Efficient Jailbreak Attack sequences on Large Language Models via Multi-Armed Bandit-based Context switching
Aditya Ramesh, Shivam Bhardwaj, Aditya Saibewar, Manohar Kaul
Abstract
Content warning: This paper contains examples of harmful language and content. Recent advances in large language models (LLMs) have made them increasingly vulnerable to jailbreaking attempts, where malicious users manipulate models into generating harmful content. While existing approaches rely on either single-step attacks that trigger immediate safety responses or multi-step methods that inefficiently iterate prompts using other LLMs, we introduce "Sequence of Context" (SoC) attacks that systematically alter conversational context through strategically crafted context-switching queries (CSQs). We formulate this as a multi-armed bandit (MAB) optimization problem, automatically learning optimal sequences of CSQs that gradually weaken the model's safety boundaries. Our theoretical analysis provides tight bounds on both the expected sequence length until successful jailbreak and the convergence of cumulative rewards. Empirically, our method achieves a 95% attack success rate, surpassing PAIR by 63.15%, AutoDAN by 60%, and ReNeLLM by 50%. We evaluate our attack across multiple open-source LLMs including Llama and Mistral variants. Our findings highlight critical vulnerabilities in current LLM safeguards and emphasize the need for defenses that consider sequential attack patterns rather than relying solely on static prompt filtering or iterative refinement. † Equal Contribution, δ was the project lead
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Red-Bandit: Test-Time Adaptation for LLM Red-Teaming via Bandit-Guided LoRA ExpertsChristos Ziakas, Nicholas Loo, Nishita Jain, Alessandra RussoACL 2026 · 2 citations
- MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative StrategiesWeiwei Qi, Shuo Shao, Wei Gu, Tianhang Zheng et al.AAAI 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
Related papers
- Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task ConcurrencyYukun Jiang, Mingjie Li, Michael Backes, Yang ZhangNeurIPS 2025 · 17 citations
- Response Attack: Exploiting Contextual Priming to Jailbreak Large Language ModelsZiqi Miao, Lijun Li, Yuan Xiong, Zhenhua Liu et al.AAAI 2026 · 8 citations
- Attention Eclipse: Manipulating Attention to Bypass LLM Safety-AlignmentPedram Zaree, Md Abdullah Al Mamun, Quazi Mishkatul Alam, Yue Dong et al.EMNLP 2025
- MASTERKEY: Automated Jailbreaking of Large Language Model ChatbotsGelei Deng, Yi Liu, Yuekang Li, Kailong Wang et al.NDSS 2024
- Multi-Turn Jailbreaking Large Language Models via Attention ShiftingXiaohu Du, Fan Mo, Ming Wen, Tu Gu et al.AAAI 2025 · 26 citations
