Persistent Semantic Entities in Tool-Augmented LLM Systems
Zhaohui Wang
摘要
Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries-largely invisible to standard debugging. We formalize this as Persistent Semantic Entities (PSE): constructs defined by name binding, event triggering, and cross-boundary propagation, and evaluate them across 24 models from 11 families (1.5B-1T parameters). First, every tested model is susceptible (20-100% on the 20model susceptibility panel), with name binding as the necessary and dominant mechanism: without it, contamination is 0%. Second, persistence depends on contamination type rather than scale or deployment: preference contamination persists undecayed on every model probed (100% at t=10) and instruction contamination persists wherever adopted, persona-style injection decays partially (90%→10%), while factual injection is model-dependent-self-corrected on Llama-3.1-8B and GPT-4o-mini but held at ceiling on both Qwen2.5-coder variants, so we do not claim it selfcorrects in general. The preference and instruction results hold across providers in our controlled setting. Third, context-isolated self-verification achieves 20-79% reduction (median 36.5%) without oracle references while keyword-based detection produces systematic false positives, and contamination compounds 1.9× along a fourstage agent pipeline (40%→75%). Preference and instruction contamination-persistent, lacking self-correction, and poorly captured by standard monitoring-represent a particularly concerning attack surface for deployed agent systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 被引用 1,715 次
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu 等ICLR 2024 · 被引用 1,469 次
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 被引用 1,085 次
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song 等NeurIPS 2024 · 被引用 539 次
相关 Paper
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious ToolsKanghua Mo, Li Hu, Yucheng Long, Zhihao LiNeurIPS 2025 · 被引用 37 次
- SSA: Semantic Contamination of LLM-Driven Fake News DetectionCheng Xu, Nan Yan, Shuhao Guan, Yuke Mei 等EMNLP 2025
- Think Twice Before You Act: Protecting LLM Agents Against Tool Description Poisoning via Isolated PlanningShanghao Shi, Xiao Wang, Chaoyu Zhang, Hao Li 等ICML 2026 · 被引用 1 次
- Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsJing Cui, Yufei Han, Jianbin Jiao, Junge ZhangAAAI 2026
- Causal Detection of Multi-Step LLM Agent AttacksViraaji Mothukuri, Reza M. PariziICML 2026 · 被引用 2 次
