In-Context Representation Hijacking
Itay Yona, Amir Sarid, Michael Karasik, Yossi Gandelsman
摘要
We introduce , a simple in-context representation hijacking attack against large language models (LLMs). The attack works by systematically replacing a harmful keyword (e.g., bomb) with a benign token (e.g., carrot) across multiple in-context examples, provided a prefix to a harmful request. We demonstrate that this substitution leads to the internal representation of the benign token converging toward that of the harmful one, effectively embedding the harmful semantics under a euphemism. As a result, superficially innocuous prompts (e.g.,"How to build a carrot?") are internally interpreted as disallowed instructions (e.g.,"How to build a bomb?"), thereby bypassing the model's safety alignment. We use interpretability tools to show that this semantic overwrite emerges layer by layer, with benign meanings in early layers converging into harmful semantics in later ones. Doublespeak is optimization-free, broadly transferable across model families, and achieves strong success rates on closed-source and open-source systems, reaching 74% ASR on Llama-3.3-70B-Instruct with a single-sentence context override. Our findings highlight a new attack surface in the latent space of LLMs, revealing that current alignment strategies are insufficient and should instead operate at the representation level.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace 等EMNLP 2020 · 被引用 1,162 次
- Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson 等NeurIPS 2024 · 被引用 835 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
相关 Paper
- Semantic Representation Attack against Aligned Large Language ModelsJiawei Lian, Jianhong Pan, Lefan Wang, Yi Wang 等NeurIPS 2025 · 被引用 3 次
- LASA: Language-Agnostic Semantic Alignment at the Semantic Bottleneck for LLM SafetyJunxiao Yang, Haoran Liu, Jinzhe Tu, Jiale Cheng 等ACL 2026 · 被引用 1 次
- Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!Zhanhui Zhou, Jie Liu, Zhichen Dong, Jiaheng Liu 等ACL 2024
- Endless Jailbreaks with Bijection LearningBrian R. Y. Huang, Maximilian Li, Leonard TangICLR 2025
- TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language ModelsZhi Xu, Jiaqi Li, Xiaotong Zhang, Hong Yu 等ICLR 2026 · 被引用 2 次
