Persistent Semantic Entities in Tool-Augmented LLM Systems
Zhaohui Wang
Abstract
Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries-largely invisible to standard debugging. We formalize this as Persistent Semantic Entities (PSE): constructs defined by name binding, event triggering, and cross-boundary propagation, and evaluate them across 24 models from 11 families (1.5B-1T parameters). First, every tested model is susceptible (20-100% on the 20model susceptibility panel), with name binding as the necessary and dominant mechanism: without it, contamination is 0%. Second, persistence depends on contamination type rather than scale or deployment: preference contamination persists undecayed on every model probed (100% at t=10) and instruction contamination persists wherever adopted, persona-style injection decays partially (90%→10%), while factual injection is model-dependent-self-corrected on Llama-3.1-8B and GPT-4o-mini but held at ceiling on both Qwen2.5-coder variants, so we do not claim it selfcorrects in general. The preference and instruction results hold across providers in our controlled setting. Third, context-isolated self-verification achieves 20-79% reduction (median 36.5%) without oracle references while keyword-based detection produces systematic false positives, and contamination compounds 1.9× along a fourstage agent pipeline (40%→75%). Preference and instruction contamination-persistent, lacking self-correction, and poorly captured by standard monitoring-represent a particularly concerning attack surface for deployed agent systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7bd0a75-c25f-4d3e-870b-d2cdbb629f49Builds on12
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Gorilla: Large Language Model Connected with Massive APIsShishir G. Patil, Tianjun Zhang, Xin Wang, Joseph E. GonzalezNeurIPS 2024 · 1,715 citations
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsYujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu et al.ICLR 2024 · 1,469 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge BasesZhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song et al.NeurIPS 2024 · 539 citations
Related papers
- Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious ToolsKanghua Mo, Li Hu, Yucheng Long, Zhihao LiNeurIPS 2025 · 37 citations
- SSA: Semantic Contamination of LLM-Driven Fake News DetectionCheng Xu, Nan Yan, Shuhao Guan, Yuke Mei et al.EMNLP 2025
- Think Twice Before You Act: Protecting LLM Agents Against Tool Description Poisoning via Isolated PlanningShanghao Shi, Xiao Wang, Chaoyu Zhang, Hao Li et al.ICML 2026 · 1 citation
- Persistent Backdoor Attacks Under Continual Fine-Tuning of LLMsJing Cui, Yufei Han, Jianbin Jiao, Junge ZhangAAAI 2026
- Causal Detection of Multi-Step LLM Agent AttacksViraaji Mothukuri, Reza M. PariziICML 2026 · 2 citations
