Many-shot Jailbreaking
Cem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Meg Tong, Jesse Mu, Daniel Ford, Francesco Mosconi, Rajashree Agrawal
摘要
We investigate a family of simple long-context attacks on large language models: prompting with hundreds of demonstrations of undesirable behavior. This is newly feasible with the larger context windows recently deployed by Anthropic, Ope-nAI and Google DeepMind. We find that in diverse, realistic circumstances, the effectiveness of this attack follows a power law, up to hundreds of shots. We demonstrate the success of this attack on the most widely used state-of-the-art closed-weight models, and across various tasks. Our results suggest very long contexts present a rich new attack surface for LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper100
- Many-Shot In-Context LearningRishabh Agarwal, Avi Singh, Lei Zhang, Bernd Bohnet 等NeurIPS 2024 · 被引用 271 次
- Rainbow Teaming: Open-Ended Generation of Diverse Adversarial PromptsMikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro 等NeurIPS 2024 · 被引用 231 次
- Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their DefensesXiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu 等NeurIPS 2024 · 被引用 96 次
- Best-of-N JailbreakingJohn Hughes, Sara Price, Aengus Lynch, Rylan Schaeffer 等NeurIPS 2025 · 被引用 78 次
- Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal TrainingYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang 等ACL 2025 · 被引用 65 次
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li 等ICLR 2024 · 被引用 419 次
- Many-Shot In-Context LearningRishabh Agarwal, Avi Singh, Lei Zhang, Bernd Bohnet 等NeurIPS 2024 · 被引用 271 次
- Steering Llama 2 via Contrastive Activation AdditionNina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong 等ACL 2024
相关 Paper
- What Really Matters in Many-Shot Attacks? An Empirical Study of Long-Context Vulnerabilities in LLMsSangyeop Kim, Yohan Lee, Yongwoo Song, Kimin LeeACL 2025
- Defenses Against Prompt Attacks Learn Surface HeuristicsShawn Li, Chenxiao Yu, Zhiyu Ni, Hao Li 等ACL 2026 · 被引用 8 次
- Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-Based Prompt Injection Attacks via the Fine-Tuning InterfaceAndrey Labunets, Nishit V. Pandya, Ashish Hooda, Xiaohan Fu 等S&P 2025
- Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking CompetitionSander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard 等EMNLP 2023 · 被引用 25 次
- Response Attack: Exploiting Contextual Priming to Jailbreak Large Language ModelsZiqi Miao, Lijun Li, Yuan Xiong, Zhenhua Liu 等AAAI 2026 · 被引用 8 次
