Many-shot Jailbreaking
Cem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Meg Tong, Jesse Mu, Daniel Ford, Francesco Mosconi, Rajashree Agrawal
Abstract
We investigate a family of simple long-context attacks on large language models: prompting with hundreds of demonstrations of undesirable behavior. This is newly feasible with the larger context windows recently deployed by Anthropic, Ope-nAI and Google DeepMind. We find that in diverse, realistic circumstances, the effectiveness of this attack follows a power law, up to hundreds of shots. We demonstrate the success of this attack on the most widely used state-of-the-art closed-weight models, and across various tasks. Our results suggest very long contexts present a rich new attack surface for LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26114eef-52c7-406a-b2f4-64b2d88a7258Cited by top-tier papers100
- Many-Shot In-Context LearningRishabh Agarwal, Avi Singh, Lei Zhang, Bernd Bohnet et al.NeurIPS 2024 · 271 citations
- Rainbow Teaming: Open-Ended Generation of Diverse Adversarial PromptsMikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro et al.NeurIPS 2024 · 231 citations
- Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their DefensesXiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu et al.NeurIPS 2024 · 96 citations
- Best-of-N JailbreakingJohn Hughes, Sara Price, Aengus Lynch, Rylan Schaeffer et al.NeurIPS 2025 · 78 citations
- Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal TrainingYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang et al.ACL 2025 · 65 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
- Many-Shot In-Context LearningRishabh Agarwal, Avi Singh, Lei Zhang, Bernd Bohnet et al.NeurIPS 2024 · 271 citations
- Steering Llama 2 via Contrastive Activation AdditionNina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong et al.ACL 2024
Related papers
- What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMsSangyeop Kim, Yohan Lee, Yongwoo Song, Kimin LeeACL 2025
- Defenses Against Prompt Attacks Learn Surface HeuristicsShawn Li, Chenxiao Yu, Zhiyu Ni, Hao Li et al.ACL 2026 · 8 citations
- Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-Based Prompt Injection Attacks via the Fine-Tuning InterfaceAndrey Labunets, Nishit V. Pandya, Ashish Hooda, Xiaohan Fu et al.S&P 2025
- Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking CompetitionSander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard et al.EMNLP 2023 · 25 citations
- Response Attack: Exploiting Contextual Priming to Jailbreak Large Language ModelsZiqi Miao, Lijun Li, Yuan Xiong, Zhenhua Liu et al.AAAI 2026 · 8 citations
