Prompt Injection as Role Confusion
Charles Ye, Jasmine Cui, Dylan Hadfield-Menell
摘要
LLMs see the world as a single stream of text, partitioned into roles like <user> or <tool>. We trace prompt injection to role confusion: models perceive the source of text from how it sounds, not its labeled role. A command hidden in a webpage hijacks an agent simply because it sounds like <user> text, despite its <tool> label. We design role probes to measure how LLMs internally perceive "who is speaking," and find that injected text occupies the same representational space as the trusted role it imitates. We demonstrate this with CoT Forgery, a zero-shot attack that injects fabricated reasoning into user prompts and tool outputs. Models mistake the forgery for their own thoughts, yielding 60% attack success against frontier models with near-zero baselines. Strikingly, the degree of role confusion predicts attack success before a single token is generated. This mechanism generalizes beyond CoT Forgery to standard agent prompt injections, revealing prompt injection as a measurable consequence of role perception. To the model, sounding like a role is indistinguishable from being one. Project page at role-confusion.github.io.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Many-shot JailbreakingCem Anil, Esin Durmus, Nina Panickssery, Mrinank Sharma 等NeurIPS 2024 · 被引用 338 次
- The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt InjectionsMilad Nasr, Nicholas Carlini, Chawin Sitawarin, Sander V. Schulhoff 等USENIX Security 2026 · 被引用 134 次
- Tensor Trust: Interpretable Prompt Injection Attacks from an Online GameSam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato 等ICLR 2024 · 被引用 123 次
相关 Paper
- The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)Zihao Wang, Yibo Jiang, Jiahao Yu, Heqing HuangICML 2025
- ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source DataReachal Wang, Yuqi Jia, Neil Zhenqiang GongNDSS 2026 · 被引用 24 次
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM AgentsHwan Chang, Yonghyun Jun, Hwanhee LeeICLR 2026 · 被引用 32 次
- Context Contamination in LLM Analysis of Network Security Logs: Poison with Passive Prompt Injection and Mitigation EvaluationRabimba Karanjai, Yang Lu, Hemanth Hegadehalli Madhavarao, Lei Xu 等USENIX Security 2026 · 被引用 4 次
- PromptLocate: Localizing Prompt Injection AttacksYuqi Jia, Yupei Liu, Zedian Shao, Jinyuan Jia 等S&P 2026 · 被引用 35 次
