Jailbreak Transferability Emerges from Shared Representations
Rico Angell, Jannik Brinkmann, He He
摘要
Jailbreak transferability is the surprising phenomenon when an adversarial attack compromising one model also elicits harmful responses from other models. Despite widespread demonstrations, there is little consensus on why transfer is possible: is it a quirk of safety training, an artifact of model families, or a more fundamental property of representation learning? We present evidence that transferability emerges from shared representations rather than incidental flaws. Across 20 open-weight models and 33 jailbreak attacks, we find two factors that systematically shape transfer: (1) representational similarity under benign prompts, and (2) the strength of the jailbreak on the source model. To move beyond correlation, we show that deliberately increasing similarity through benign-only distillation causally increases transfer. Our qualitative analyses reveal systematic transferability patterns across different types of jailbreaks. For example, persona-style jailbreaks transfer far more often than cipher-based prompts, consistent with the idea that natural-language attacks exploit models' shared representation space, whereas cipher-based attacks rely on idiosyncratic quirks that do not generalize. Together, these results reframe jailbreak transfer as a consequence of representation alignment rather than a fragile byproduct of safety training. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Furina: Fragmented Uncertainty-Driven Refusal Instability AttackTongxi Wu, Jian Zhang, Yang GaoICML 2026 · 被引用 1 次
- One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMsYixin Tan, Yu Zhe, Rui Wen, Jun SakumaCCS 2026
- Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering AttractorsYen-Shan Chen, Sian-Yao Huang, Cheng-Lin Yang, Yun-Nung ChenICML 2026
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen 等ICLR 2024 · 被引用 1,104 次
相关 Paper
- Enhancing the Transferability of Jailbreak Attacks on Large Language Models via Exploiting Reparameterization InvarianceAo Wang, Xinghao Yang, Yongshun Gong, Wei Liu 等ACL 2026
- Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous ConstraintsJunxiao Yang, Zhexin Zhang, Shiyao Cui, Hongning Wang 等ACL 2025 · 被引用 6 次
- Towards Understanding Jailbreak Attacks in LLMs: A Representation Space AnalysisYuping Lin, Pengfei He, Han Xu, Yue Xing 等EMNLP 2024 · 被引用 6 次
- Attention Eclipse: Manipulating Attention to Bypass LLM Safety-AlignmentPedram Zaree, Md Abdullah Al Mamun, Quazi Mishkatul Alam, Yue Dong 等EMNLP 2025
- Failures to Find Transferable Image Jailbreaks Between Vision-Language ModelsRylan Schaeffer, Dan Valentine, Luke Bailey, James Chua 等ICLR 2025
