ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
Hwan Chang, Yonghyun Jun, Hwanhee Lee
摘要
The growing deployment of large language model (LLM) based agents that interact with external environments has created new attack surfaces for adversarial manipulation. One major threat is indirect prompt injection, where attackers embed malicious instructions in external environment output, causing agents to interpret and execute them as if they were legitimate prompts. While previous research has focused primarily on plain-text injection attacks, we find a significant yet underexplored vulnerability: LLMs' dependence on structured chat templates and their susceptibility to contextual manipulation through persuasive multi-turn dialogues. To this end, we introduce ChatInject, an attack that formats malicious payloads to mimic native chat templates, thereby exploiting the model's inherent instruction-following tendencies. Building on this foundation, we develop a template-based Multi-turn variant that primes the agent across conversational turns to accept and execute otherwise suspicious actions. Through comprehensive experiments across frontier LLMs, we demonstrate three critical findings: (1) ChatInject achieves significantly higher average attack success rates than traditional prompt injection methods, improving from 5.18% to 32.05% on AgentDojo and from 15.13% to 45.90% on InjecAgent, with multi-turn dialogues showing particularly strong performance at average 52.33% success rate on InjecAgent, (2) chat-template-based payloads demonstrate strong transferability across models and remain effective even against closed-source LLMs, despite their unknown template structures, and (3) existing prompt-based defenses are largely ineffective against this attack approach, especially against Multi-turn variants. These findings highlight vulnerabilities in current agent systems. The code is available at https://hwanchang00.github.io/chatinject_project_page .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 被引用 1,026 次
- Bad Characters: Imperceptible NLP AttacksNicholas Boucher, Ilia Shumailov, Ross Anderson, Nicolas PapernotS&P 2022 · 被引用 133 次
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMsYi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang 等ACL 2024 · 被引用 64 次
- On the Loss of Context Awareness in General Instruction Fine-tuningYihan Wang, Andrew Bai, Nanyun Peng, Cho-Jui HsiehNeurIPS 2025 · 被引用 11 次
- Foot-In-The-Door: A Multi-turn Jailbreak for LLMsZixuan Weng, Xiaolong Jin, Jinyuan Jia, Xiangyu ZhangEMNLP 2025 · 被引用 2 次
相关 Paper
- It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web AgentsKarolina Korgul, Yushi Yang, Arkadiusz Drohomirecki, Piotr Blaszczyk 等ICML 2026 · 被引用 8 次
- TopicAttack: An Indirect Prompt Injection Attack via Topic TransitionYulin Chen, Haoran Li, Yuexin Li, Yue Liu 等EMNLP 2025 · 被引用 1 次
- AgentBreaker: Evaluating Context-Aware Indirect Prompt Injection Risks in Modern Web AgentsYongbi Son, Changoo Lee, Dongwon Shin, Byoungyoung Lee 等ISSTA 2026
- A Large-scale Measurement of In-Page Prompt Injections Against LLM Web AgentsSoheil Khodayari, Xuenan Zhang, Bhupendra Acharya, Giancarlo PellegrinoCCS 2026
- Manipulating Multimodal Agents via Cross-Modal Prompt InjectionLe Wang, Zonghao Ying, Tianyuan Zhang, Siyuan Liang 等ACM MM 2025 · 被引用 8 次
